dt-evals is an open source CLI tool from Dynatrace that enables teams to evaluate LLM and AI agent output quality by pulling real GenAI traces, scoring them with an LLM-as-judge approach, and writing structured results back into Dynatrace AI Observability. It supports both offline (pre-release CI/CD) and online (post-deployment sampling) evaluation modes. Built-in evaluators cover 15 quality and safety dimensions including faithfulness, hallucination, relevance, toxicity, PII leakage, and prompt injection, with support for custom evaluators. Results integrate with Dynatrace dashboards, DQL queries, and alerting workflows, enabling teams to trend quality scores, detect regressions, and gate releases on AI quality metrics alongside traditional latency and error rate signals.

13m read timeFrom dynatrace.com
Post cover image
Table of contents
What is dt-evals?What are LLM evaluations?Run evaluations from the command lineBring your own LLM judge providerWhich quality and safety dimensions are evaluated by dt-evals?Evaluation results in the AI Observability appQuery, trend, and alert on evaluation scoresHow to turn quality regressions into alertsClose the loop in the AI software delivery lifecycleBring evaluations into the release processComing nextStart today
180 Impressions