Aphrodite Engine
High-throughput inference engine for large language models
Aphrodite Engine is the official backend serving engine for PygmalionAI, optimized for high-throughput LLM inference. It exposes an OpenAI-compatible API and supports a wide range of quantization formats.
Key features
- High-throughput batching
- OpenAI-compatible endpoints
- Many quantization formats
- Continuous batching
Strengths
- Released under the AGPL-3.0 license
- First-class Docker support for quick deployment
- Active community (1.8k GitHub stars)
- Written in C++
Aphrodite Engine replaces
Compare Aphrodite Engine
25 head-to-head comparisons.
- Aphrodite Engine vs Ollama
- Aphrodite Engine vs llama.cpp
- Aphrodite Engine vs vLLM
- Aphrodite Engine vs GPT4All
- Aphrodite Engine vs LiteLLM
- Aphrodite Engine vs exo
- Aphrodite Engine vs New API
- Aphrodite Engine vs Jan
- Aphrodite Engine vs FastChat
- Aphrodite Engine vs One API
- Aphrodite Engine vs Continue
- Aphrodite Engine vs llamafile
- Aphrodite Engine vs MLC LLM
- Aphrodite Engine vs OpenLLM
- Aphrodite Engine vs KoboldCpp
- Aphrodite Engine vs Text Generation Inference
- Aphrodite Engine vs Petals
- Aphrodite Engine vs Page Assist
- Aphrodite Engine vs LMDeploy
- Aphrodite Engine vs Enchanted
- Aphrodite Engine vs Serge
- Aphrodite Engine vs LM Studio
- Aphrodite Engine vs Text Embeddings Inference
- Aphrodite Engine vs Harbor LLM Toolkit
- Aphrodite Engine vs ik_llama.cpp
Similar local llm runners apps
Ollama
Local LLM RunnersRun large language models locally with a simple CLI and API
Replaces ChatGPT, OpenAI API
llama.cpp
Local LLM RunnersHigh-performance LLM inference in plain C/C++
Replaces OpenAI API
vLLM
Local LLM RunnersHigh-throughput LLM serving engine with PagedAttention
Replaces OpenAI API
GPT4All
Local LLM RunnersPrivacy-first desktop chat with local language models
Replaces ChatGPT
LiteLLM
Local LLM RunnersUnified proxy and gateway for 100+ LLM APIs
Replaces OpenRouter
exo
Local LLM RunnersRun your own AI cluster across everyday devices
Replaces OpenAI API