LMDeploy
Toolkit for compressing and serving large language models
LMDeploy is an inference and serving toolkit from the OpenMMLab ecosystem for deploying large language models efficiently. It offers quantization, a high-throughput serving engine, and an OpenAI-compatible API.
Key features
- High-throughput inference engine
- Weight quantization support
- OpenAI-compatible serving
- Multi-GPU deployment
Strengths
- Released under the Apache-2.0 license
- First-class Docker support for quick deployment
- Mature project with 8.1k GitHub stars
- Written in Python
LMDeploy replaces
Compare LMDeploy
15 head-to-head comparisons.
- LMDeploy vs Ollama
- LMDeploy vs llama.cpp
- LMDeploy vs vLLM
- LMDeploy vs New API
- LMDeploy vs exo
- LMDeploy vs FastChat
- LMDeploy vs One API
- LMDeploy vs llamafile
- LMDeploy vs MLC LLM
- LMDeploy vs OpenLLM
- LMDeploy vs Text Generation Inference
- LMDeploy vs Petals
- LMDeploy vs ik_llama.cpp
- LMDeploy vs Aphrodite Engine
- LMDeploy vs Wllama
Similar local llm runners apps
Ollama
Local LLM RunnersRun large language models locally with a simple CLI and API
Replaces ChatGPT, OpenAI API
llama.cpp
Local LLM RunnersHigh-performance LLM inference in plain C/C++
Replaces OpenAI API
vLLM
Local LLM RunnersHigh-throughput LLM serving engine with PagedAttention
Replaces OpenAI API
GPT4All
Local LLM RunnersPrivacy-first desktop chat with local language models
Replaces ChatGPT
LiteLLM
Local LLM RunnersUnified proxy and gateway for 100+ LLM APIs
Replaces OpenRouter
New API
Local LLM RunnersNext-gen LLM gateway and AI asset management system
Replaces OpenRouter, OpenAI API