vLLM
High-throughput LLM serving engine with PagedAttention
vLLM is a fast and memory-efficient inference and serving engine for large language models. Its PagedAttention algorithm delivers high throughput batching, and it exposes an OpenAI-compatible server for production deployments.
Key features
- PagedAttention memory management
- Continuous batching
- OpenAI-compatible server
- Tensor parallelism
Strengths
- Released under the Apache-2.0 license
- First-class Docker support for quick deployment
- Kubernetes-ready with Helm charts available
- Mature project with 88.5k GitHub stars
vLLM replaces
Compare vLLM
36 head-to-head comparisons.
- vLLM vs Ollama
- vLLM vs Hugging Face Transformers
- vLLM vs llama.cpp
- vLLM vs GPT4All
- vLLM vs GPT4Free
- vLLM vs LiteLLM
- vLLM vs LocalAI
- vLLM vs exo
- vLLM vs New API
- vLLM vs Jan
- vLLM vs FastChat
- vLLM vs One API
- vLLM vs Continue
- vLLM vs SGLang
- vLLM vs llamafile
- vLLM vs MLC LLM
- vLLM vs Guidance
- vLLM vs OpenLLM
- vLLM vs KoboldCpp
- vLLM vs Text Generation Inference
- vLLM vs Petals
- vLLM vs Xinference
- vLLM vs Llama Stack
- vLLM vs Page Assist
- vLLM vs LMDeploy
- vLLM vs MLX LM
- vLLM vs Enchanted
- vLLM vs Serge
- vLLM vs GPUStack
- vLLM vs LM Studio
- vLLM vs Text Embeddings Inference
- vLLM vs Harbor LLM Toolkit
- vLLM vs ik_llama.cpp
- vLLM vs Aphrodite Engine
- vLLM vs Wllama
- vLLM vs LLMKube
Similar local llm runners apps
Ollama
Local LLM RunnersRun large language models locally with a simple CLI and API
Replaces ChatGPT, OpenAI API
llama.cpp
Local LLM RunnersHigh-performance LLM inference in plain C/C++
Replaces OpenAI API
GPT4All
Local LLM RunnersPrivacy-first desktop chat with local language models
Replaces ChatGPT
LiteLLM
Local LLM RunnersUnified proxy and gateway for 100+ LLM APIs
Replaces OpenRouter
exo
Local LLM RunnersRun your own AI cluster across everyday devices
Replaces OpenAI API
New API
Local LLM RunnersNext-gen LLM gateway and AI asset management system
Replaces OpenRouter, OpenAI API