OpenLLM
Run any open LLM as an OpenAI-compatible API endpoint
OpenLLM lets developers run open-source large language models as OpenAI-compatible API endpoints with a single command. It is built on BentoML and supports streaming, quantization, and easy deployment.
Key features
- One-command model serving
- OpenAI-compatible API
- Built-in chat UI
- Cloud deployment via BentoML
Strengths
- Released under the Apache-2.0 license
- First-class Docker support for quick deployment
- Kubernetes-ready with Helm charts available
- Mature project with 12.5k GitHub stars
OpenLLM replaces
Compare OpenLLM
15 head-to-head comparisons.
Similar local llm runners apps
Ollama
Local LLM RunnersRun large language models locally with a simple CLI and API
Replaces ChatGPT, OpenAI API
llama.cpp
Local LLM RunnersHigh-performance LLM inference in plain C/C++
Replaces OpenAI API
vLLM
Local LLM RunnersHigh-throughput LLM serving engine with PagedAttention
Replaces OpenAI API
GPT4All
Local LLM RunnersPrivacy-first desktop chat with local language models
Replaces ChatGPT
LiteLLM
Local LLM RunnersUnified proxy and gateway for 100+ LLM APIs
Replaces OpenRouter
New API
Local LLM RunnersNext-gen LLM gateway and AI asset management system
Replaces OpenRouter, OpenAI API