DE

DeepEval

Unit testing framework for LLM outputs

Self-Hosted AI ★ 17.5k stars Medium setup Apache-2.0

DeepEval is an open-source evaluation framework that brings unit-testing-style assertions to large language model outputs. It includes metrics for hallucination, relevancy, and bias and integrates with pytest for local CI.

Key features

  • Pytest-style LLM tests
  • 14+ built-in metrics
  • Synthetic dataset generation
  • CI/CD integration

Strengths

  • Released under the Apache-2.0 license
  • Mature project with 17.5k GitHub stars
  • Written in Python

DeepEval replaces

Compare DeepEval

18 head-to-head comparisons.

Similar self-hosted ai apps