Google's Gemini Enterprise Agent Platform now offers generally available agent and model evaluations, enabling consistent quality measurement from development through production. The system supports offline experiments against curated test cases and online monitoring of live traffic using the same metrics engine. Key features include 20+ pre-built metrics (Task Success, Tool Use Quality, Safety, Hallucination, Grounding), adaptive LLM-judge rubrics co-developed with Google DeepMind, custom code-based and prompt-based metrics, and automated test case generation via case generators, user simulators, and environment simulators. Evaluations integrate with the Agent Platform SDK, agents-cli, ADK, and the Google Cloud console, with artifacts stored in Cloud Storage for auditability.