llama.cpp
High-performance LLM inference in plain C/C++
llama.cpp is a C/C++ inference engine for running LLaMA-family and many other models efficiently on CPUs and GPUs. It pioneered the GGUF quantized model format and powers a large portion of the local-AI ecosystem.
Key features
- GGUF quantization
- Runs on modest hardware
- Built-in HTTP server
- Broad GPU backend support
Strengths
- Released under the MIT license
- First-class Docker support for quick deployment
- Mature project with 123k GitHub stars
- Written in C++
llama.cpp replaces
Compare llama.cpp
36 head-to-head comparisons.
- llama.cpp vs Ollama
- llama.cpp vs Hugging Face Transformers
- llama.cpp vs vLLM
- llama.cpp vs GPT4All
- llama.cpp vs GPT4Free
- llama.cpp vs LiteLLM
- llama.cpp vs LocalAI
- llama.cpp vs exo
- llama.cpp vs New API
- llama.cpp vs Jan
- llama.cpp vs FastChat
- llama.cpp vs One API
- llama.cpp vs Continue
- llama.cpp vs SGLang
- llama.cpp vs llamafile
- llama.cpp vs MLC LLM
- llama.cpp vs Guidance
- llama.cpp vs OpenLLM
- llama.cpp vs KoboldCpp
- llama.cpp vs Text Generation Inference
- llama.cpp vs Petals
- llama.cpp vs Xinference
- llama.cpp vs Llama Stack
- llama.cpp vs Page Assist
- llama.cpp vs LMDeploy
- llama.cpp vs MLX LM
- llama.cpp vs Enchanted
- llama.cpp vs Serge
- llama.cpp vs GPUStack
- llama.cpp vs LM Studio
- llama.cpp vs Text Embeddings Inference
- llama.cpp vs Harbor LLM Toolkit
- llama.cpp vs ik_llama.cpp
- llama.cpp vs Aphrodite Engine
- llama.cpp vs Wllama
- llama.cpp vs LLMKube
Similar local llm runners apps
Ollama
Local LLM RunnersRun large language models locally with a simple CLI and API
Replaces ChatGPT, OpenAI API
vLLM
Local LLM RunnersHigh-throughput LLM serving engine with PagedAttention
Replaces OpenAI API
GPT4All
Local LLM RunnersPrivacy-first desktop chat with local language models
Replaces ChatGPT
LiteLLM
Local LLM RunnersUnified proxy and gateway for 100+ LLM APIs
Replaces OpenRouter
exo
Local LLM RunnersRun your own AI cluster across everyday devices
Replaces OpenAI API
New API
Local LLM RunnersNext-gen LLM gateway and AI asset management system
Replaces OpenRouter, OpenAI API