OpenLLM
Run any open LLM as an OpenAI-compatible API endpoint
OpenLLM lets developers run open-source large language models as OpenAI-compatible API endpoints with a single command. It is built on BentoML and supports streaming, quantization, and easy deployment.
Key features
- One-command model serving
- OpenAI-compatible API
- Built-in chat UI
- Cloud deployment via BentoML
Strengths
- Released under the Apache-2.0 license
- First-class Docker support for quick deployment
- Kubernetes-ready with Helm charts available
- Mature project with 12.5k GitHub stars
OpenLLM replaces
Compare OpenLLM
34 head-to-head comparisons.
- OpenLLM vs Ollama
- OpenLLM vs Hugging Face Transformers
- OpenLLM vs llama.cpp
- OpenLLM vs vLLM
- OpenLLM vs GPT4All
- OpenLLM vs GPT4Free
- OpenLLM vs LiteLLM
- OpenLLM vs LocalAI
- OpenLLM vs exo
- OpenLLM vs New API
- OpenLLM vs Jan
- OpenLLM vs FastChat
- OpenLLM vs One API
- OpenLLM vs Continue
- OpenLLM vs SGLang
- OpenLLM vs llamafile
- OpenLLM vs MLC LLM
- OpenLLM vs Guidance
- OpenLLM vs KoboldCpp
- OpenLLM vs Text Generation Inference
- OpenLLM vs Petals
- OpenLLM vs Xinference
- OpenLLM vs Llama Stack
- OpenLLM vs Page Assist
- OpenLLM vs LMDeploy
- OpenLLM vs MLX LM
- OpenLLM vs Enchanted
- OpenLLM vs Serge
- OpenLLM vs GPUStack
- OpenLLM vs LM Studio
- OpenLLM vs Text Embeddings Inference
- OpenLLM vs Harbor LLM Toolkit
- OpenLLM vs ik_llama.cpp
- OpenLLM vs Aphrodite Engine
Similar local llm runners apps
Ollama
Local LLM RunnersRun large language models locally with a simple CLI and API
Replaces ChatGPT, OpenAI API
llama.cpp
Local LLM RunnersHigh-performance LLM inference in plain C/C++
Replaces OpenAI API
vLLM
Local LLM RunnersHigh-throughput LLM serving engine with PagedAttention
Replaces OpenAI API
GPT4All
Local LLM RunnersPrivacy-first desktop chat with local language models
Replaces ChatGPT
LiteLLM
Local LLM RunnersUnified proxy and gateway for 100+ LLM APIs
Replaces OpenRouter
exo
Local LLM RunnersRun your own AI cluster across everyday devices
Replaces OpenAI API