A Microsoft and open source AI lab Hugging Face benchmark grades 507 multi step business workflows — tasks whose later steps depend on what earlier steps changed in a system of record — on what the database actually says at the end, exposing a
The agent industry has been grading the wrong thing. Most benchmarks score an AI agent on whether it made the right tool calls, not whether the work behind those calls actually got done. A new joint Microsoft and Hugging Face benchmark called ThinkingBox flips that, and exposes a failure class the old scoring could not see.
A grader that checks tool calls sees nine well-formed calls and marks the run a success. A grader that checks the underlying record at the end sees those same nine calls left the work unfinished. ThinkingBox uses the second grader. It checks terminal backend state and side effects, the same way a human auditor would walk over to the system of record and ask what actually changed. The same nine calls can look like a triumph on one grader and a regression on the other.
A buyer in Nashville is waiting on a $745 kitchen appliance. The package sits in a courier "exception" for fifteen days past the estimated delivery date. The agent handles the case by the policy: it pulls up the order, contacts the carrier, files the appropriate follow-up, and closes the ticket. The closing line is "Since your query is resolved, is there anything I may assist you with?" The benchmark marks this run a failure. The ticket status ends as "solved." The required end state is "hold." The carrier exception is still open.
That is the gap ThinkingBox is built to measure. A standard evaluator looking at the tool trace would record nine successful calls and a clean policy. ThinkingBox records the wrong terminal state. The run that "passed" every tool-call check still ships a customer back to a courier queue with no escalation, because the carrier exception is still open. The case lives in the benchmark as test_case_ST003_006, with the full failure trace in Appendix D.4, Case 3 of the underlying paper.
The harness is wider than one anecdote. ThinkingBox covers 507 stateful business workflows, where a stateful workflow is a task whose later steps depend on what earlier steps changed in the system of record. Each workflow runs 20 times against each model under test. That repetition is the point: one successful run is not reliability. A model that fails one in twenty attempts is shipping a production bug at scale, and the 507-by-20 matrix is the smallest statistical base the benchmark is willing to call stable.
The failure shapes repeat across the harness. Some agents over-close, ending on a "solved" status the policy forbids. Some agents under-write, recording the right intent but never flipping the underlying flag from "exception" to "hold" or "refund_pending." Some agents pick the right tool but the wrong record, so the change lands on a different ticket. Each signature looks correct in a tool trace. None of them leave the right terminal state.
Two pieces of plumbing are worth knowing on first encounter. MCP is a standard way for an AI agent to call external tools, the contract that lets an agent talk to a ticket system, a CRM, or a shipping API without a custom integration per product. OpenEnv is Hugging Face's runnable sandbox harness; the ThinkingBox tasks and their graders are executable through it today, which means a team shipping agents can drop the benchmark into a continuous integration loop and catch this failure class before it reaches a customer.
The release is joint work. The named co-authors on the blog and paper are Tommy Guy, founder at Enderis AI and previously Microsoft; Sergio Paniego at Hugging Face; Zhuochun Li at the University of Pittsburgh; Ali Keramati at UC Irvine; and Youngmin Ko at Northwestern. The blog is a primary source from both labs; the methodology, the 507-workflow harness, and the OpenEnv runnable artifact are all described there.
The next agent shipping in production will be graded by something. Today, most production graders look at the tool trace. ThinkingBox looks at the database, and it is runnable today.