LiteLLM
Unified proxy and gateway for 100+ LLM APIs
LiteLLM provides a proxy server and Python SDK that exposes a single OpenAI-compatible interface to over a hundred LLM providers. It adds load balancing, spend tracking, caching, and key management.
Key features
- One API for many providers
- Spend and rate tracking
- Load balancing and fallbacks
- Virtual API keys
Strengths
- Released under the MIT license
- First-class Docker support for quick deployment
- Kubernetes-ready with Helm charts available
- Mature project with 55.8k GitHub stars
LiteLLM replaces
Compare LiteLLM
27 head-to-head comparisons.
- LiteLLM vs Ollama
- LiteLLM vs llama.cpp
- LiteLLM vs vLLM
- LiteLLM vs GPT4All
- LiteLLM vs exo
- LiteLLM vs New API
- LiteLLM vs Jan
- LiteLLM vs FastChat
- LiteLLM vs One API
- LiteLLM vs Continue
- LiteLLM vs llamafile
- LiteLLM vs MLC LLM
- LiteLLM vs OpenLLM
- LiteLLM vs KoboldCpp
- LiteLLM vs Text Generation Inference
- LiteLLM vs Petals
- LiteLLM vs Page Assist
- LiteLLM vs LMDeploy
- LiteLLM vs Enchanted
- LiteLLM vs Serge
- LiteLLM vs LM Studio
- LiteLLM vs Text Embeddings Inference
- LiteLLM vs Harbor LLM Toolkit
- LiteLLM vs ik_llama.cpp
- LiteLLM vs Aphrodite Engine
- LiteLLM vs Wllama
- LiteLLM vs LLMKube
Similar local llm runners apps
Ollama
Local LLM RunnersRun large language models locally with a simple CLI and API
Replaces ChatGPT, OpenAI API
llama.cpp
Local LLM RunnersHigh-performance LLM inference in plain C/C++
Replaces OpenAI API
vLLM
Local LLM RunnersHigh-throughput LLM serving engine with PagedAttention
Replaces OpenAI API
GPT4All
Local LLM RunnersPrivacy-first desktop chat with local language models
Replaces ChatGPT
exo
Local LLM RunnersRun your own AI cluster across everyday devices
Replaces OpenAI API
New API
Local LLM RunnersNext-gen LLM gateway and AI asset management system
Replaces OpenRouter, OpenAI API