Ollama
Local LLM RunnersRun large language models locally with a simple CLI and API
Replaces ChatGPT, OpenAI API
Self-hosted local large-language-model inference engines and serving frameworks.
28 self-hosted apps · 345 comparisons
Run large language models locally with a simple CLI and API
Replaces ChatGPT, OpenAI API
High-performance LLM inference in plain C/C++
Replaces OpenAI API
High-throughput LLM serving engine with PagedAttention
Replaces OpenAI API
Privacy-first desktop chat with local language models
Replaces ChatGPT
Unified proxy and gateway for 100+ LLM APIs
Replaces OpenRouter
Run your own AI cluster across everyday devices
Replaces OpenAI API
Next-gen LLM gateway and AI asset management system
Replaces OpenRouter, OpenAI API
Open-source offline ChatGPT alternative for the desktop
Replaces ChatGPT
Platform for serving and evaluating large language models
Replaces OpenAI API
Unified OpenAI-compatible gateway for many LLM providers
Replaces OpenAI API, OpenRouter
Open-source AI code assistant for VS Code and JetBrains
Replaces GitHub Copilot, Cursor
Distribute and run LLMs with a single executable file
Replaces OpenAI API, ChatGPT
Universal LLM deployment engine for any hardware
Replaces OpenAI API
Run any open LLM as an OpenAI-compatible API endpoint
Replaces OpenAI API
Single-file local LLM runner for text and storytelling
Replaces ChatGPT
Hugging Face toolkit for production LLM serving
Replaces OpenAI API
Run large language models collaboratively in a swarm
Replaces OpenAI API
Browser extension to use local AI models on the web
Replaces ChatGPT
Toolkit for compressing and serving large language models
Replaces OpenAI API, Hugging Face Inference Endpoints
Native Ollama client for iOS and macOS
Replaces ChatGPT
Self-hosted LLaMA chat UI with no API keys needed
Replaces ChatGPT
Desktop app to discover, download, and run local LLMs
Replaces ChatGPT
Fast inference server for text embedding models
Replaces OpenAI Embeddings API
Containerized LLM toolkit to run a local AI stack with one CLI
Replaces OpenAI Platform
Performance-focused fork of llama.cpp with new quant types
Replaces OpenAI API
High-throughput inference engine for large language models
Replaces OpenAI API
Run LLM inference directly in the browser with WebAssembly
Replaces OpenAI API
Kubernetes operator for llama.cpp-native LLM inference with GPU
No apps match these filters.
345 head-to-head comparisons in this category.