Databricks AI Research Team has released OfficeQA Pro V2, a new benchmark for evaluating AI agents on enterprise grounded-reasoning tasks. Built from approximately 120,000 pages of U.S. Treasury Accounts of Receipts and Expenditures spanning 1793–2024, the benchmark contains 90 questions requiring document retrieval, analytical reasoning, and multimodal interpretation across a previously unseen corpus. Out-of-the-box frontier agents (Claude Code and Codex) averaged 37.5% accuracy, while competition teams averaged 41.1% and the winner reached 63.3%. Databricks' own Genie agent improved accuracy by an average of 24 percentage points over baseline harnesses. The benchmark is publicly available on Hugging Face and GitHub, and is intended to test whether AI agent improvements generalize beyond a single document corpus — a key concern in real enterprise deployments.