LM

LM Evaluation Harness

Unified framework to benchmark language models on many tasks

Self-Hosted AI ★ 14.1k stars Medium setup MIT

The Language Model Evaluation Harness from EleutherAI is a framework for evaluating generative language models on a large number of standardized benchmarks. It can be run locally to assess self-hosted models.

Key features

  • Hundreds of benchmarks
  • Many model backends
  • Reproducible evaluation
  • Standard in research

Pros & cons

Strengths

  • De facto research standard
  • Hundreds of tasks included
  • Supports local model backends

Trade-offs

  • Not a server or web app
  • Evaluations can take hours

LM Evaluation Harness replaces

Similar self-hosted ai apps