<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/the-agent-said-it-was-done-the-database-disagreed--th0ihs2vh" -->

---
title: The Agent Said It Was Done. The Database Disagreed.
description: Microsoft's ThinkingBox, now available on Hugging Face via OpenEnv, benchmarks AI agents by checking the actual terminal database state and side effects they...
canonical: https://daily.dev/posts/the-agent-said-it-was-done-the-database-disagreed--th0ihs2vh
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: The Agent Said It Was Done. The Database Disagreed. | daily.dev
og:description: Microsoft's ThinkingBox, now available on Hugging Face via OpenEnv, benchmarks AI agents by checking the actual terminal database state and side effects they...
og:url: https://daily.dev/posts/the-agent-said-it-was-done-the-database-disagreed--th0ihs2vh
og:image: https://api.daily.dev/og/posts/TH0iHS2Vh.png
og:image:alt: The Agent Said It Was Done. The Database Disagreed.
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# The Agent Said It Was Done. The Database Disagreed.

**[Hugging Face](https://daily.dev/sources/huggingface)** · 16 min read · 2 upvotes · 0 comments

## Summary

Microsoft's ThinkingBox, now available on Hugging Face via OpenEnv, benchmarks AI agents by checking the actual terminal database state and side effects they leave behind rather than trusting their final responses or tool calls. Each of 507 stateful business workflows (retail, auto insurance, travel, neobank, consulting) is run 20 times per model across 18 LLMs to separate single-attempt capability from true reliability. Across a 12-model ablation, 67% of failed attempts still looked clean (no tool errors), yet executable checks found wrong field values in 77.6% of those failures. Claude Opus 5.5 leads overall pass@1 at 67.16%, while Kimi-K3 has the broadest one-shot coverage (93.89%) but poor consistency (13.41% pass 20/20 times). Cost analysis shows the cheapest model per single success isn't the cheapest per dependable task, and roughly 80% of failures trace to tool-handling/error-recovery problems rather than reasoning. The benchmark, dataset, and harness are open for anyone to reproduce.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://huggingface.co/blog/microsoft/thinkingbox>

## Questions this post answers

### Which LLM is most reliable for AI agents that need to consistently complete the same task every time, not just once?

Claude Opus 5 and Claude Opus 5.5 are the most consistent, each completing 47.53% of 507 business-workflow tasks correctly on all 20 repeated attempts, versus only 13.41% for Kimi-K3 despite Kimi-K3 solving more tasks at least once (93.89%). Pass@1 or pass@20 scores can mask this gap between occasional success and true dependability.

_Developers choosing a model for production agents can track benchmark comparisons like this one on daily.dev._

### Why does an AI agent report a task as resolved when the underlying database record is still incomplete or incorrect?

Agents can produce a clean final response and valid tool calls while still leaving wrong field values, missing required effects, or unintended extra side effects in the backend. In one large ablation, 67.24% of failed attempts terminated cleanly with no tool error, yet 77.61% of those failures had wrong field values and 43.30% had unintended extra effects, since checking tool calls alone does not verify the resulting database state.

_Anyone building agent evaluation pipelines can follow discussions on state-based grading via daily.dev._

### What causes most AI agent task failures in enterprise workflows, reasoning errors or tool usage problems?

About 79.9% of failures stem from tool usage issues rather than reasoning failures, based on a deterministic failure-signature analysis of agent traces. Other causes include wrong state updates (10.3%), incomplete user resolutions (7.0%), and no state-changing action taken (2.9%), suggesting retry logic and error recovery around tool calls matter more than model reasoning quality for fixing these failures.

_Teams debugging flaky agent behavior can compare failure patterns and fixes through daily.dev._

## Similar posts on daily.dev

- [The Agentic Analytics Benchmark: Measuring model accuracy and efficiency in analytical agents](https://daily.dev/posts/the-agentic-analytics-benchmark-measuring-model-accuracy-and-efficiency-in-analytical-agents-81a3conc5) · ClickHouse · 0 upvotes · 0 comments
- [Your Agent Aced the Task. Will It Do It Again?](https://daily.dev/posts/your-agent-aced-the-task-will-it-do-it-again--yetkhtezb) · Hugging Face · 0 upvotes · 0 comments
- [Claude did best on a new benchmark for agents that build agents. It still passed fewer than a quarter of the tests.](https://daily.dev/posts/claude-did-best-on-a-new-benchmark-for-agents-that-build-agents-it-still-passed-fewer-than-a-quarte-rrx7ttpl4) · The New Stack · 2 upvotes · 0 comments
- [Introducing o11y-bench: an open benchmark for AI agents running observability workflows](https://daily.dev/posts/introducing-o11y-bench-an-open-benchmark-for-ai-agents-running-observability-workflows-psucsa908) · Grafana Labs · 20 upvotes · 0 comments

---

Tags: [#llm](https://daily.dev/tags/llm), [#ai-agents](https://daily.dev/tags/ai-agents), [#mcp](https://daily.dev/tags/mcp)

[View this post on daily.dev](https://daily.dev/posts/the-agent-said-it-was-done-the-database-disagreed--th0ihs2vh)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"The Agent Said It Was Done. The Database Disagreed.","url":"https://daily.dev/posts/the-agent-said-it-was-done-the-database-disagreed--th0ihs2vh","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/the-agent-said-it-was-done-the-database-disagreed--th0ihs2vh"},"datePublished":"2026-10-03T22:49:20.884Z","dateModified":"2026-10-03T22:49:47.279Z","description":"Microsoft's ThinkingBox, now available on Hugging Face via OpenEnv, benchmarks AI agents by checking the actual terminal database state and side effects they...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/69d67b6791b578eccddd7726d15b4d7f?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/69d67b6791b578eccddd7726d15b4d7f?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"Hugging Face","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Hugging Face","logo":"https://media.daily.dev/image/upload/t_logo,f_auto/v1/logos/f1f55c67d81a4330acf5b90b26b0c8e1","url":"https://daily.dev/sources/huggingface"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/the-agent-said-it-was-done-the-database-disagreed--th0ihs2vh","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":2},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"llm,ai-agents,mcp","timeRequired":"PT16M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Hugging Face","item":"https://daily.dev/sources/huggingface"},{"@type":"ListItem","position":3,"name":"The Agent Said It Was Done. The Database Disagreed."}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/the-agent-said-it-was-done-the-database-disagreed--th0ihs2vh#faq","mainEntity":[{"@type":"Question","name":"Which LLM is most reliable for AI agents that need to consistently complete the same task every time, not just once?","acceptedAnswer":{"@type":"Answer","text":"Claude Opus 5 and Claude Opus 5.5 are the most consistent, each completing 47.53% of 507 business-workflow tasks correctly on all 20 repeated attempts, versus only 13.41% for Kimi-K3 despite Kimi-K3 solving more tasks at least once (93.89%). Pass@1 or pass@20 scores can mask this gap between occasional success and true dependability. Developers choosing a model for production agents can track benchmark comparisons like this one on daily.dev."}},{"@type":"Question","name":"Why does an AI agent report a task as resolved when the underlying database record is still incomplete or incorrect?","acceptedAnswer":{"@type":"Answer","text":"Agents can produce a clean final response and valid tool calls while still leaving wrong field values, missing required effects, or unintended extra side effects in the backend. In one large ablation, 67.24% of failed attempts terminated cleanly with no tool error, yet 77.61% of those failures had wrong field values and 43.30% had unintended extra effects, since checking tool calls alone does not verify the resulting database state. Anyone building agent evaluation pipelines can follow discussions on state-based grading via daily.dev."}},{"@type":"Question","name":"What causes most AI agent task failures in enterprise workflows, reasoning errors or tool usage problems?","acceptedAnswer":{"@type":"Answer","text":"About 79.9% of failures stem from tool usage issues rather than reasoning failures, based on a deterministic failure-signature analysis of agent traces. Other causes include wrong state updates (10.3%), incomplete user resolutions (7.0%), and no state-changing action taken (2.9%), suggesting retry logic and error recovery around tool calls matter more than model reasoning quality for fixing these failures. Teams debugging flaky agent behavior can compare failure patterns and fixes through daily.dev."}}]}
```

