EX

ExLlamaV2

Fast inference library for quantized LLMs on consumer GPUs

Self-Hosted AI ★ 4.6k stars Hard setup MIT

ExLlamaV2 is an inference library optimized for running quantized large language models efficiently on modern consumer GPUs. Its EXL2 quantization format allows flexible bitrates for the best speed-quality balance.

Key features

  • EXL2 flexible quantization
  • Fast single-GPU inference
  • Low memory footprint
  • Built-in server

Strengths

  • Released under the MIT license
  • Active community (4.6k GitHub stars)
  • Written in Python

ExLlamaV2 replaces

Compare ExLlamaV2

10 head-to-head comparisons.

Similar self-hosted ai apps