LM

LMDeploy

Toolkit for compressing and serving large language models

Local LLM Runners ★ 8k stars Hard setup Apache-2.0

LMDeploy is an inference and serving toolkit from the OpenMMLab ecosystem for deploying large language models efficiently. It offers quantization, a high-throughput serving engine, and an OpenAI-compatible API.

Key features

  • High-throughput inference engine
  • Weight quantization support
  • OpenAI-compatible serving
  • Multi-GPU deployment

Strengths

  • Released under the Apache-2.0 license
  • First-class Docker support for quick deployment
  • Mature project with 8k GitHub stars
  • Written in Python

LMDeploy replaces

Compare LMDeploy

34 head-to-head comparisons.

Similar local llm runners apps