Aphrodite Engine
High-throughput inference engine for large language models
Aphrodite Engine is the official backend serving engine for PygmalionAI, optimized for high-throughput LLM inference. It exposes an OpenAI-compatible API and supports a wide range of quantization formats.
Key features
- High-throughput batching
- OpenAI-compatible endpoints
- Many quantization formats
- Continuous batching
Strengths
- Released under the AGPL-3.0 license
- First-class Docker support for quick deployment
- Active community (1.9k GitHub stars)
- Written in Python
Aphrodite Engine replaces
Compare Aphrodite Engine
15 head-to-head comparisons.
- Aphrodite Engine vs Ollama
- Aphrodite Engine vs llama.cpp
- Aphrodite Engine vs vLLM
- Aphrodite Engine vs New API
- Aphrodite Engine vs exo
- Aphrodite Engine vs FastChat
- Aphrodite Engine vs One API
- Aphrodite Engine vs llamafile
- Aphrodite Engine vs MLC LLM
- Aphrodite Engine vs OpenLLM
- Aphrodite Engine vs Text Generation Inference
- Aphrodite Engine vs Petals
- Aphrodite Engine vs LMDeploy
- Aphrodite Engine vs ik_llama.cpp
- Aphrodite Engine vs Wllama
Similar local llm runners apps
Ollama
Local LLM RunnersRun large language models locally with a simple CLI and API
Replaces ChatGPT, OpenAI API
llama.cpp
Local LLM RunnersHigh-performance LLM inference in plain C/C++
Replaces OpenAI API
vLLM
Local LLM RunnersHigh-throughput LLM serving engine with PagedAttention
Replaces OpenAI API
GPT4All
Local LLM RunnersPrivacy-first desktop chat with local language models
Replaces ChatGPT
LiteLLM
Local LLM RunnersUnified proxy and gateway for 100+ LLM APIs
Replaces OpenRouter
New API
Local LLM RunnersNext-gen LLM gateway and AI asset management system
Replaces OpenRouter, OpenAI API