Datadog Agent Observability now supports running DeepEval and Pydantic Evals evaluation frameworks natively within Datadog Experiments, eliminating the need to rewrite existing evaluators or adopt proprietary metric definitions. The integration lets teams define datasets, configure existing evaluators without modification, and run experiments that automatically link eval scores to production traces, token usage, and latency data. Teams can also run evaluations continuously on sampled production traffic rather than only as a pre-deployment CI gate, enabling real-time quality regression detection across the full development and deployment lifecycle.

6m read timeFrom datadoghq.com
Post cover image
Table of contents
Why framework portability matters for LLM evalsSet up experiments with Datadog Agent ObservabilityAnalyze experiment results in DatadogConnect eval scores to production tracesRun LLM evals continuously on production trafficGet started with Datadog Agent Observability
453 Impressions