EX

ExLlama

Memory-efficient inference library for quantized Llama models

Self-Hosted AI ★ 2.9k stars Hard setup MIT

ExLlama is a standalone Python/C++/CUDA implementation for running quantized GPTQ Llama models with low VRAM use on modern GPUs. It is the predecessor to ExLlamaV2 and focuses on fast, memory-efficient local inference.

Key features

  • Low VRAM GPTQ inference
  • CUDA-accelerated
  • Standalone library
  • Fast token generation

Strengths

  • Released under the MIT license
  • Active community (2.9k GitHub stars)
  • Written in Python

ExLlama replaces

Compare ExLlama

10 head-to-head comparisons.

Similar self-hosted ai apps