---
title: "EVA-Bench Data 2.0: 3 Domains, 121 Tools, 213 Scenarios"
url: https://daily.dev/posts/eva-bench-data-2-0-3-domains-121-tools-213-scenarios-8h0ks6ujs
source_url: https://huggingface.co/blog/ServiceNow-AI/eva-bench-data
type: article
source: "Hugging Face"
published: 2026-06-04T12:29:09.916Z
updated: 2026-06-04T12:29:34.319Z
tags: ["llm"]
reading_time: 10
upvotes: 0
comments: 0
language: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# EVA-Bench Data 2.0: 3 Domains, 121 Tools, 213 Scenarios

**[Hugging Face](https://daily.dev/sources/huggingface)** · 10 min read · 0 upvotes · 0 comments

## Summary

EVA-Bench 2.0 is an open-source benchmark for evaluating enterprise voice agents, expanding from one domain to three: Airline Customer Service Management, Enterprise IT Service Management, and Healthcare HR Service Delivery. The release covers 213 evaluation scenarios across 121 tools — roughly 4x the original coverage. Scenarios are generated using SyGra, a graph-based synthetic data pipeline, and validated against three frontier models (GPT-5.4, Gemini 3.1 Pro, Claude Opus 4.6). Key design principles include voice-first scope, realism, scenario variety (single-intent, multi-intent, adversarial), authentication flows, and strict reproducibility via deterministic user goals and ground-truth database states. A multilingual extension is also previewed. All datasets are available on Hugging Face under the MIT license.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://huggingface.co/blog/ServiceNow-AI/eva-bench-data>

## Similar posts on daily.dev

- [A New Framework for Evaluation of Voice Agents \(EVA\)](https://daily.dev/posts/a-new-framework-for-evaluation-of-voice-agents-eva--k34p62hv2) · Hugging Face · 0 upvotes · 0 comments
- [ITBench-AA: Frontier Models Score Below 50% on the First Benchmark for Agentic Enterprise IT Tasks — by Artificial Analysis and IBM](https://daily.dev/posts/itbench-aa-frontier-models-score-below-50-on-the-first-benchmark-for-agentic-enterprise-it-tasks--inm3wceim) · Hugging Face · 1 upvotes · 0 comments
- [Enterprise AI benchmarks are broken](https://daily.dev/posts/enterprise-ai-benchmarks-are-broken-os7hvrdov) · The New Stack · 0 upvotes · 0 comments

---

Tags: [#llm](https://daily.dev/tags/llm)

[View this post on daily.dev](https://daily.dev/posts/eva-bench-data-2-0-3-domains-121-tools-213-scenarios-8h0ks6ujs)
