Saturday, October 3, 2026 Newsletter Advertise
AI

Microsoft ThinkingBox grades AI agents on database state, not text

Microsoft and Hugging Face released a benchmark that runs 507 enterprise agent workflows 20 times each and scores the records agents leave behind.

Microsoft ThinkingBox grades AI agents on database state, not text. Source: Hugging Face

Microsoft and Hugging Face published ThinkingBox on October 3, 2026, an agent sandbox and benchmark that scores AI agents on the backend state and side effects they leave behind rather than on their tool calls or final messages. The companies said the harness and dataset are now available through Hugging Face and the OpenEnv interface.

What the benchmark measures

According to the joint blog post by Microsoft and Hugging Face, ThinkingBox is the agent sandbox and ThinkingBox-Bench is the dataset used to evaluate agents. The authors wrote that across 507 stateful business workflows, each run 20 times against various LLM models, it grades agents on terminal backend state and side effects.

Each task defines a starting backend state, a user goal, the available MCP tools, the domain policy, and executable checks over the terminal state, the post said, with a simulated user that holds private context and releases it only when asked. Every attempt gets an isolated MCP session with freshly initialized state, and 477 of the 507 tasks are graded on state alone while 30 add response rubrics.

The post opens with a retail example in which an agent makes nine well-formed tool calls, then closes a ticket as solved when the required end state is hold. The authors said the failing check is a single field.

Clean-looking runs that still fail

In a common-set ablation covering 121,680 valid trials across 12 LLM models, 79,853 attempts failed the executable checks, the authors reported. Of those failures, 67.24% still terminated cleanly, invoked a state-changing tool, and reported no final tool error.

Executable checks nevertheless found wrong field values in 77.61% of them, unintended extra effects in 43.30%, and missing required effects in 25.36%, with those findings overlapping, according to the post.

Microsoft and Hugging Face also assigned each failed trace one deterministic diagnostic signature. Tool usage accounted for 79.9% of failures, wrong state updates 10.3%, incomplete user resolutions 7.0% and no state-changing action 2.9%. The authors described this as a retry and error-recovery problem before it is a model problem.

Breadth versus consistency

The post reports pass@1, pass@20 and an observed 20/20 count. Claude Opus 5.5 leads overall at 67.16% pass@1, the authors said, while Kimi-K3 is described as the strongest open-weights model at 57.37% overall and 82.24% on retail.

Kimi-K3 solves 93.89% of the benchmark at least once, or 476 of 507 tasks, but only 68 tasks, 13.41%, succeed in all 20 attempts, according to the post. Claude Opus 5 solves fewer tasks at least once, 79.09%, but completes 47.53% of the benchmark on every attempt.

Claude Opus 5.5 passes exactly the same number of tasks on all 20 attempts as Claude Opus 5: 241. The authors wrote that if you are choosing a model for work that touches real records, pass@20 is the wrong column to look at. Cost figures in the post are estimates priced at undiscounted list rates on OpenRouter from a snapshot taken on September 20, 2026, and the authors describe them as a comparative efficiency index, not an invoice.

Availability and caveats

ThinkingBox-Bench now sits behind the OpenEnv interface, and each finished episode returns a binary pass/fail reward, the post said. The released adapter is designed for evaluation, with separate non-benchmark scenarios usable for training workflows.

ThinkingBox code is MIT-licensed, the benchmark data is CDLA-Permissive-2.0 and the OpenEnv environment ships under BSD-3-Clause. The authors noted that every task in the public benchmark is a synthetic reconstruction and the customers are not real.

Microsoft said ThinkingBox and ThinkingBox-Bench is built by the Microsoft Copilot Studio team in partnership with Toloka, with collaborators from the University of Pittsburgh, Northwestern University, Columbia University and UC Irvine who interned at Microsoft.

What to do

  • Treat the 20/20 rate as a design input rather than a verdict when assessing agent reliability, the authors say.
  • Check the terminal state of your backend before committing a change, instead of trusting the model's summary of what it did.
  • Classify tool and system errors so that retries target the recoverable ones.
  • Reduce the tool surface exposed to an agent to only what the workflow needs.
  • Require human approval for changes that cannot be cheaply reversed.
  • The authors note they have not measured the benefit of these steps on the benchmark, and say the environment now makes that testable.
  • To reproduce results, the post says you need Linux or WSL, Python 3.11+, uv and Docker, a thinkingbox-data checkout at the pinned release, and model endpoints for the agent, simulated user and judge; one endpoint can serve all three roles.
  • Report a repeat metric and define it, stating whether it is best-of-k or every-of-k and how it was computed.
Key facts and where they come from
  • The benchmark covers 507 stateful business workflows, each run 20 times, graded on terminal backend state and side effects.
    Across 507 stateful business workflows, each run 20 times against various LLM models, it grades agents on terminal backend state and side effects.
  • In an ablation of 121,680 valid trials across 12 models, 79,853 attempts failed the executable checks.
    In a common-set ablation covering 121,680 valid trials across 12 LLM models, 79,853 attempts failed the executable checks.
  • Most failures looked clean: 67.24% terminated cleanly, used a state-changing tool and reported no final tool error.
    Of those failures, 67.24% still terminated cleanly, invoked a state-changing tool, and reported no final tool error.
  • Claude Opus 5.5 posted the highest overall pass@1 at 67.16%.
    Claude Opus 5.5 leads overall at 67.16%, two-thirds of a point above Claude Opus 5.
  • Kimi-K3 solved 476 of 507 tasks at least once but passed all 20 attempts on only 68 tasks.
    It solves 93.89% of the benchmark at least once: 476 of 507 tasks.
  • Tool usage accounts for 79.9% of failure signatures.
    roughly four in five failures are tool handling, not reasoning
  • Most tasks are graded purely on state; a minority add response rubrics.
    477 of the 507 tasks are graded on state alone; 30 add response rubrics.
  • The harness and dataset are published on Hugging Face.
    ThinkingBox is now on Hugging Face, both the harness and the dataset.
  • All benchmark tasks are synthetic reconstructions.
    every task in the public benchmark is a synthetic reconstruction
  • Code and data use permissive licenses.
    ThinkingBox code is MIT-licensed; the benchmark data is CDLA-Permissive-2.0; the OpenEnv environment ships under OpenEnv's BSD-3-Clause.

Read the original from Hugging Face →

The TechUpscale Brief

The day's cyber, AI and tech news in one short email, every weekday morning. Free. Unsubscribe anytime.

I'm most interested in

More AI