MLC LLM
Universal LLM deployment engine for any hardware
MLC LLM is a machine learning compiler and runtime that deploys language models natively across GPUs, CPUs, browsers, and mobile devices. It enables high-performance self-hosted inference on diverse hardware.
Key features
- Compile models for any hardware
- Native GPU acceleration
- Browser and mobile runtimes
- OpenAI-compatible serving
Strengths
- Released under the Apache-2.0 license
- Mature project with 23k GitHub stars
- Written in Python
MLC LLM replaces
Compare MLC LLM
34 head-to-head comparisons.
- MLC LLM vs Ollama
- MLC LLM vs Hugging Face Transformers
- MLC LLM vs llama.cpp
- MLC LLM vs vLLM
- MLC LLM vs GPT4All
- MLC LLM vs GPT4Free
- MLC LLM vs LiteLLM
- MLC LLM vs LocalAI
- MLC LLM vs exo
- MLC LLM vs New API
- MLC LLM vs Jan
- MLC LLM vs FastChat
- MLC LLM vs One API
- MLC LLM vs Continue
- MLC LLM vs SGLang
- MLC LLM vs llamafile
- MLC LLM vs Guidance
- MLC LLM vs OpenLLM
- MLC LLM vs KoboldCpp
- MLC LLM vs Text Generation Inference
- MLC LLM vs Petals
- MLC LLM vs Xinference
- MLC LLM vs Llama Stack
- MLC LLM vs Page Assist
- MLC LLM vs LMDeploy
- MLC LLM vs MLX LM
- MLC LLM vs Enchanted
- MLC LLM vs Serge
- MLC LLM vs GPUStack
- MLC LLM vs LM Studio
- MLC LLM vs Text Embeddings Inference
- MLC LLM vs Harbor LLM Toolkit
- MLC LLM vs ik_llama.cpp
- MLC LLM vs Aphrodite Engine
Similar local llm runners apps
Ollama
Local LLM RunnersRun large language models locally with a simple CLI and API
Replaces ChatGPT, OpenAI API
llama.cpp
Local LLM RunnersHigh-performance LLM inference in plain C/C++
Replaces OpenAI API
vLLM
Local LLM RunnersHigh-throughput LLM serving engine with PagedAttention
Replaces OpenAI API
GPT4All
Local LLM RunnersPrivacy-first desktop chat with local language models
Replaces ChatGPT
LiteLLM
Local LLM RunnersUnified proxy and gateway for 100+ LLM APIs
Replaces OpenRouter
exo
Local LLM RunnersRun your own AI cluster across everyday devices
Replaces OpenAI API