Ollama
Local LLM RunnersRun large language models locally with a simple CLI and API
Replaces ChatGPT, OpenAI API
The 14 best Docker local llm runners you can self-host, ranked by community traction.
Run large language models locally with a simple CLI and API
Replaces ChatGPT, OpenAI API
High-throughput LLM serving engine with PagedAttention
Replaces OpenAI API
Unified proxy and gateway for 100+ LLM APIs
Replaces OpenRouter
Next-gen LLM gateway and AI asset management system
Replaces OpenRouter, OpenAI API
Platform for serving and evaluating large language models
Replaces OpenAI API
Unified OpenAI-compatible gateway for many LLM providers
Replaces OpenAI API, OpenRouter
Run any open LLM as an OpenAI-compatible API endpoint
Replaces OpenAI API
Hugging Face toolkit for production LLM serving
Replaces OpenAI API
Toolkit for compressing and serving large language models
Replaces OpenAI API, Hugging Face Inference Endpoints
Self-hosted LLaMA chat UI with no API keys needed
Replaces ChatGPT
Fast inference server for text embedding models
Replaces OpenAI Embeddings API
Containerized LLM toolkit to run a local AI stack with one CLI
Replaces OpenAI Platform
High-throughput inference engine for large language models
Replaces OpenAI API
Kubernetes operator for llama.cpp-native LLM inference with GPU
No apps match these filters.
Every option here is open-source and self-hostable, tagged Docker within the local llm runners category. Compare them on the individual app pages for setup difficulty and resource requirements.