LM

LMDeploy

Toolkit for compressing and serving large language models

Local LLM Runners ★ 8.1k stars Hard setup Apache-2.0

LMDeploy is an inference and serving toolkit from the OpenMMLab ecosystem for deploying large language models efficiently. It offers quantization, a high-throughput serving engine, and an OpenAI-compatible API.

Key features

  • High-throughput inference engine
  • Weight quantization support
  • OpenAI-compatible serving
  • Multi-GPU deployment

Strengths

  • Released under the Apache-2.0 license
  • First-class Docker support for quick deployment
  • Mature project with 8.1k GitHub stars
  • Written in Python

LMDeploy replaces

Compare LMDeploy

15 head-to-head comparisons.

Similar local llm runners apps