Nine correct-looking calls, one wrong outcome

Microsoft and Hugging Face describe a customer-support agent that checked a delayed $745 appliance order, opened a ticket and read the policy correctly—then marked the unresolved case solved. The tool trace looked competent. The database still had the wrong final state.

ThinkingBox evaluates the record and side effects the agent leaves behind. Its benchmark covers 507 synthetic workflows in retail, insurance, travel, banking and consulting. Every task is run 20 times from a clean backend, so the results distinguish one successful attempt from repeated reliability.

Why execution traces can mislead

Across 121,680 trials for 12 models, 79,853 attempts failed executable checks. Microsoft reports that 67.24% of those failures still ended cleanly after a state-changing call, with no final tool error. The evaluator found wrong field values in 77.61% of failed trials, unintended extra effects in 43.30%, and missing required effects in 25.36%; those categories overlap.

The lesson is practical: a polished answer or valid tool call is only a proxy for success. For agents touching business records, the check should inspect the resulting state and whether the requested side effects actually happened.

A benchmark builders can run

ThinkingBox and its dataset are available through Hugging Face’s OpenEnv adapter. The release is designed to let teams evaluate their own model against isolated tool sessions and deterministic checks. Most tasks are graded directly on state, with a smaller set of response rubrics for requirements that are not represented in a database field.

The benchmark is synthetic, so it does not establish how a deployed support agent performs on real customers. Its value is the test design: repeat workflows, capture terminal state, and expose the quiet failures that ordinary trace review can miss.

Explore the original source ↗

Source published October 3, 2026. Coverage is based on the maker’s announcement and demonstration.