VL

vLLM

High-throughput LLM serving engine with PagedAttention

Local LLM Runners ★ 88.5k stars Hard setup Apache-2.0

vLLM is a fast and memory-efficient inference and serving engine for large language models. Its PagedAttention algorithm delivers high throughput batching, and it exposes an OpenAI-compatible server for production deployments.

Key features

  • PagedAttention memory management
  • Continuous batching
  • OpenAI-compatible server
  • Tensor parallelism

Strengths

  • Released under the Apache-2.0 license
  • First-class Docker support for quick deployment
  • Kubernetes-ready with Helm charts available
  • Mature project with 88.5k GitHub stars

vLLM replaces

Compare vLLM

36 head-to-head comparisons.

Similar local llm runners apps