dt-evals is an open source CLI tool from Dynatrace that enables teams to evaluate LLM and AI agent output quality by pulling real GenAI traces, scoring them with an LLM-as-judge approach, and writing structured results back into Dynatrace AI Observability. It supports both offline (pre-release CI/CD) and online (post-deployment sampling) evaluation modes. Built-in evaluators cover 15 quality and safety dimensions including faithfulness, hallucination, relevance, toxicity, PII leakage, and prompt injection, with support for custom evaluators. Results integrate with Dynatrace dashboards, DQL queries, and alerting workflows, enabling teams to trend quality scores, detect regressions, and gate releases on AI quality metrics alongside traditional latency and error rate signals.