databricks
Read post

Introducing OfficeQA Pro V2: A New Benchmark for Enterprise Grounded-Reasoning

Databricks AI Research Team has released OfficeQA Pro V2, a new benchmark for evaluating AI agents on enterprise grounded-reasoning tasks. Built from approximately 120,000 pages of U.S. Treasury Accounts of Receipts and Expenditures spanning 1793–2024, the benchmark contains 90 questions requiring document retrieval, analytical reasoning, and multimodal interpretation across a previously unseen corpus. Out-of-the-box frontier agents (Claude Code and Codex) averaged 37.5% accuracy, while competition teams averaged 41.1% and the winner reached 63.3%. Databricks' own Genie agent improved accuracy by an average of 24 percentage points over baseline harnesses. The benchmark is publicly available on Hugging Face and GitHub, and is intended to test whether AI agent improvements generalize beyond a single document corpus — a key concern in real enterprise deployments.

    #ai-agents#rag#databricks
Yesterday•12m read time•From databricks.com
Post cover image
Table of contents
Agent Performance on OfficeQA Pro V2Building a New Grounded Reasoning BenchmarkScaling OfficeQA Pro V2 with Synthetic DataExample QuestionsBenchmark DetailsConclusion & Acknowledgements
21 Impressions
databricks's image
databricks

461 Followers

•

1.4K Upvotes

Would you recommend this post?

Copy link
WhatsApp
Facebook
X
New Squad
  • © 2026 Daily Dev Ltd.
  • Guidelines
  • Explore
  • Tags
  • Sources
  • Squads
  • Leaderboard