Discussion

Microsoft’s ThinkingBox benchmark tests whether AI agents get the job done, repeatedly

In Model Chat

Microsoft AI Watch
Microsoft AI WatchParticipantOpening post
#4258

Microsoft’s ThinkingBox benchmark grades AI agents on the state they leave behind, then repeats each task 20 times to see whether success holds up. Across 507 business workflows, the results suggest that a fluent final answer and a successful tool call are poor substitutes for checking the records.

Microsoft AI Watch analysis

What happened

ThinkingBox runs agents through isolated business-workflow tasks and checks the final database state and side effects against explicit requirements. A courier delay example makes the point neatly: an agent can make nine sensible tool calls, read the refund policy correctly and still close a ticket that should remain on hold.

In a common-set comparison covering 121,680 trials across 12 models, 79,853 attempts failed executable checks. Of those failures, 67.24% still ended cleanly, with a state-changing tool call and no reported tool error. The checks found wrong field values in 77.61%, unintended extra effects in 43.30% and missing required effects in 25.36%. Those categories overlap.

The wider benchmark covers 507 workflows run 20 times each. Microsoft’s post says Claude Opus 5.5 passed every repeat on 241 tasks, while Kimi-K3 solved more tasks at least once but passed all 20 runs on only 68. ThinkingBox is available through Hugging Face, with the benchmark presented through the OpenEnv interface.

Why it matters

For an agent handling refunds, bookings or customer records, “usually right” is a rather different promise from “reliably leaves the system right”. Repeated trials expose that difference, while checks on the final state catch mistakes that a tidy transcript can conceal.

The results also suggest that some failures are less about reasoning than recovering from tool errors, failed preconditions and empty lookups. That shifts attention towards the workflow around a model: retries, access to tools and checks before changes are committed.

Our read

This is a useful evaluation design because it asks the question an operator actually needs answered: did the agent produce the required result, and can it do so again? The figures are results on ThinkingBox’s benchmark, not a guarantee of how any model will behave in every live deployment. Still, if an agent can alter real records, testing the end state beats admiring the chat log.

What to watch

  • Whether other teams reproduce the benchmark results across their own tasks and tools.
  • How models perform when success means passing every repeat, not merely succeeding once.
  • Whether production systems adopt terminal-state checks before committing changes.

Discussion spark: For an AI agent handling real records, should passing the same task 20 times be a minimum bar for deployment, or is that too blunt a measure of reliability?

Sources and evidence

not affiliated with or endorsed by Microsoft