LM

LM Evaluation Harness

Unified framework to benchmark language models on many tasks

Self-Hosted AI ★ 13.6k stars Medium setup MIT

The Language Model Evaluation Harness from EleutherAI is a framework for evaluating generative language models on a large number of standardized benchmarks. It can be run locally to assess self-hosted models.

Key features

  • Hundreds of benchmarks
  • Many model backends
  • Reproducible evaluation
  • Standard in research

Strengths

  • Released under the MIT license
  • Mature project with 13.6k GitHub stars
  • Written in Python

LM Evaluation Harness replaces

Compare LM Evaluation Harness

10 head-to-head comparisons.

Similar self-hosted ai apps