ExLlamaV2
Fast inference library for quantized LLMs on consumer GPUs
ExLlamaV2 is an inference library optimized for running quantized large language models efficiently on modern consumer GPUs. Its EXL2 quantization format allows flexible bitrates for the best speed-quality balance.
Key features
- EXL2 flexible quantization
- Fast single-GPU inference
- Low memory footprint
- Built-in server
Strengths
- Released under the MIT license
- Active community (4.6k GitHub stars)
- Written in Python
ExLlamaV2 replaces
Compare ExLlamaV2
10 head-to-head comparisons.
Similar self-hosted ai apps
OpenClaw
Self-Hosted AIThe AI that actually does things
Hermes Agent
Self-Hosted AIThe AI agent that grows with you
OpenCode
Self-Hosted AIThe open source AI coding agent
Hugging Face Transformers
Self-Hosted AIState-of-the-art machine learning model library
Replaces OpenAI API
Langflow
Self-Hosted AIVisual framework for building AI agents and RAG pipelines
Replaces Vertex AI Agent Builder
Dify
Self-Hosted AIOpen-source platform for building production LLM apps
Replaces OpenAI Assistants, Vertex AI Agent Builder