DeepEval
Unit testing framework for LLM outputs
DeepEval is an open-source evaluation framework that brings unit-testing-style assertions to large language model outputs. It includes metrics for hallucination, relevancy, and bias and integrates with pytest for local CI.
Key features
- Pytest-style LLM tests
- 14+ built-in metrics
- Synthetic dataset generation
- CI/CD integration
Strengths
- Released under the Apache-2.0 license
- Mature project with 17.5k GitHub stars
- Written in Python
DeepEval replaces
Compare DeepEval
18 head-to-head comparisons.
- DeepEval vs OpenClaw
- DeepEval vs Hermes Agent
- DeepEval vs OpenCode
- DeepEval vs Hugging Face Transformers
- DeepEval vs Langflow
- DeepEval vs Dify
- DeepEval vs PaddleOCR
- DeepEval vs RAGFlow
- DeepEval vs OpenHands
- DeepEval vs screenshot-to-code
- DeepEval vs Langfuse
- DeepEval vs promptfoo
- DeepEval vs Ragas
- DeepEval vs Prompt flow
- DeepEval vs Arize Phoenix
- DeepEval vs OpenLLMetry
- DeepEval vs Agenta
- DeepEval vs Weave
Similar self-hosted ai apps
OpenClaw
Self-Hosted AIThe AI that actually does things
Hermes Agent
Self-Hosted AIThe AI agent that grows with you
OpenCode
Self-Hosted AIThe open source AI coding agent
Hugging Face Transformers
Self-Hosted AIState-of-the-art machine learning model library
Replaces OpenAI API
Langflow
Self-Hosted AIVisual framework for building AI agents and RAG pipelines
Replaces Vertex AI Agent Builder
Dify
Self-Hosted AIOpen-source platform for building production LLM apps
Replaces OpenAI Assistants, Vertex AI Agent Builder