EVA-Bench 2.0 is an open-source benchmark for evaluating enterprise voice agents, expanding from one domain to three: Airline Customer Service Management, Enterprise IT Service Management, and Healthcare HR Service Delivery. The release covers 213 evaluation scenarios across 121 tools — roughly 4x the original coverage. Scenarios are generated using SyGra, a graph-based synthetic data pipeline, and validated against three frontier models (GPT-5.4, Gemini 3.1 Pro, Claude Opus 4.6). Key design principles include voice-first scope, realism, scenario variety (single-intent, multi-intent, adversarial), authentication flows, and strict reproducibility via deterministic user goals and ground-truth database states. A multilingual extension is also previewed. All datasets are available on Hugging Face under the MIT license.

10m read timeFrom huggingface.co
Post cover image
Table of contents
IntroductionData Design PrinciplesScenario GenerationFurther ValidationDataset Deep-DivesMultilingual SupportGet the DataCitations
142 Impressions