The Agent Said It Was Done. The Database Counted 241.

4 minute read

Published:

中文版 Chinese version

Microsoft and Hugging Face released ThinkingBox on October 3, 2026. It runs 507 stateful business workflows — retail, auto insurance, travel, neobank, consulting — 20 times per model from an identical clean backend, and grades the terminal database state and side effects, not the transcript. 477 tasks are graded on state alone; 30 add response rubrics.

That design choice exposes a gap most leaderboards hide.

The headline tie

Claude Opus 5.5 leads pass@1 at 67.16%. Claude Opus 5 is at 66.50%. A +0.66 point gain.

On observed 20/20 — tasks that passed all 20 recorded attempts, no estimator, no smoothing — both models pass exactly 241 of 507 tasks (47.53%).

Half a point of headline accuracy bought zero additional dependable tasks.

Breadth pulls the other way. Kimi-K3, the strongest open-weight model in the set at 57.37% pass@1, solves 476 of 507 tasks at least once (93.89%). Only 68 tasks (13.41%) pass all 20 times. Opus 5 solves fewer tasks at least once (79.09%) but holds 241 of them every time — 173 more consistently-passed tasks than Kimi-K3.

If you pick on pass@20, you pick the flaky model. If you pick on pass@1, you still cannot tell whether the gain survives repetition. GPT-6 Astra retains 78% of its single-attempt rate across repeats; Opus 5.5 and Opus 5 each retain 71%. GLM-5.1, Kimi-K2.6 and DeepSeek-V4-Pro each keep about 8%.

ThinkingBox: pass@1 vs passing all 20 runs, and cost per dependable task

A clean finish is not evidence

In the common-set ablation — 121,680 valid trials across 12 models — 79,853 attempts failed the executable checks. Of those failures, 67.24% terminated cleanly, invoked a state-changing tool, and reported no final tool error.

The checks found wrong field values in 77.61% of failures, unintended extra effects in 43.30%, and missing required effects in 25.36% (findings overlap).

The diagnostic signature is more useful than the score. Tool handling accounts for 79.9% of failures (unweighted average of per-model shares). Wrong state updates are 10.3%, incomplete user resolutions 7.0%, and no state-changing action 2.9%. Agents usually attempt the workflow, then fail to recover from a tool error, an unmet precondition, or an empty lookup. That is a retry and error-recovery problem before it is a model problem.

Domain spread reinforces the point: retail averages 59.52% pass@1 across models; auto insurance averages 33.83%.

Price dependability, not single successes

The authors price recorded token usage at undiscounted list rates (a comparative index, not an invoice) two ways.

Cost per successful attempt rewards cheap, often-right answers. GPT-5.6 Sol is lowest there at $0.127 per success.

Cost per dependable task — the full 20-run campaign cost divided by tasks passing 20/20 — reorders the field:

ModelTasks passing 20/20Cost per dependable task
GPT-5.4128$6.80
GPT-6 Astra231$7.45
Claude Opus 5.5241$7.80
GPT-5.6 Sol82$9.76
Claude Opus 5241$13.30
Kimi-K368$20.68

Opus 5.5 dominates Opus 5 on this measure: same 241 dependable tasks, $7.80 vs $13.30. GPT-5.6 Sol, cheapest per single success, costs $9.76 per dependable task. The cheapest way to get a right answer is not the cheapest way to get a dependable one.

Caveats

These are the benchmark authors’ own runs and list-rate estimates, not independent replication or your cloud bill. Observed 20/20 records what happened in these trials; it is not a guarantee of future runs. The authors state they have not measured the lift from their own suggested fixes — terminal-state checks, error classification, human approval on irreversible changes. Treat those as testable hypotheses, not proven gains.

What to do Monday

  1. Score your own replay on 20/20, not pass@1. Take 30–50 of your production workflows, run each 20 times from a clean backend, and grade final state. If observed 20/20 is below 40% on a workflow that writes irreversible state, keep human approval on that workflow until it clears 40% on two consecutive weekly replays.
  2. Add two columns to every model comparison: tasks passing all 20 runs, and cost per dependable task (campaign cost / 20/20 tasks). If a newer model gains under 1 point of pass@1 and adds zero 20/20 tasks — the exact Opus 5 to 5.5 pattern — do not pay a migration cost for dependability you did not get.
  3. Budget the retry layer before the next model. With 79.9% of failures in tool handling, spend the first sprint on error classification and targeted retries for recoverable tool errors, and re-measure. Only escalate the model if 20/20 is still below your threshold after that.

A trajectory is a claim. Database state is the evidence. Repetition is the trust test — and now it has a price.

Sources: Microsoft / Hugging Face, “The Agent Said It Was Done. The Database Disagreed.”, October 3, 2026.