SG

SGLang

Fast serving framework for LLMs and vision-language models

Self-Hosted AI ★ 31.5k stars Hard setup Apache-2.0

SGLang is a high-performance serving framework for large language and vision-language models. It features a fast runtime with RadixAttention and a flexible programming language for complex LLM applications.

Key features

  • RadixAttention caching
  • Structured generation
  • OpenAI-compatible server
  • Multi-GPU scaling

Strengths

  • Released under the Apache-2.0 license
  • First-class Docker support for quick deployment
  • Kubernetes-ready with Helm charts available
  • Mature project with 31.5k GitHub stars

SGLang replaces

Compare SGLang

30 head-to-head comparisons.

Similar self-hosted ai apps