Practical AI guides

Evaluation & debugging

Choose checks that expose regressions, inspect agent behavior, and make failures easier to diagnose.

Where to start

Separate answer-quality evaluation from operational tracing and security scanning. Each answers a different question. Begin with a representative failure you want to detect, then examine coverage and blind spots.

Explore the reports

EvaluationpromptfooSep 53 min read

Promptfoo: building evaluation checks that catch real regressions

24.8Kstars
EvaluationlangfuseSep 23 min read

Langfuse: turning an LLM trace into a useful debugging decision

34.1Kstars
EvaluationTencentAug 233 min read

AI-Infra-Guard: choosing the right scan for your AI stack

5.7Kstars
Search all reports in this topic