surehand
All articlesOperations

Microsoft graded agents on the database, not the reply. Your approver should too.

News4 min readSurehand

When an agent says "done", that is a claim, not a fact. Microsoft just put a number on how often the claim is wrong, and it is high enough that you should change what your approver looks at. If your approver reads the agent's summary, they are grading the wrong thing.

On 3 October, Microsoft's Copilot Studio team and Hugging Face released ThinkingBox, a test bench for agents doing real business work1. It has 507 workflows across retail, auto insurance, travel, a neobank and consulting. Each one runs 20 times, from a clean database every time. The agent is graded on what it left in the database, not on what it said.

The tasks are synthetic, modelled on real enterprise patterns, with made-up customers1. The failure modes are not made up. You will recognise them.

Nine good tool calls. One wrong field.

Here is the example Microsoft leads with1. A customer's $745 kitchen appliance is stuck in a courier "exception" at a Nashville distribution center, fifteen days late. The agent does careful work. Nine tool calls. It pulls the order, checks tracking, reads the refund policy correctly and opens a ticket.

Then it closes the ticket as resolved.

The courier exception was still open. The right end state was "on hold". The customer never got a real answer either. Anyone checking the tool calls saw nine well-formed ones. Only the ticket's status field showed the failure.

Now picture your version. Your AP agent marks an invoice "paid" while the payment run sits in a failed batch. Your support agent closes the claim before the refund posts. The trace looks fine. The record is wrong.

Most failures look like success

The headline numbers come from 121,680 valid runs across 12 models1. 79,853 of them failed the checks. Of those failures, 67.24% still ended cleanly, called a tool that changed data, and reported no final tool error.

So about two in three failed runs would pass a glance at the log.

What was actually wrong in those runs? Wrong field values in 77.61%. Extra changes nobody asked for in 43.30%. Required changes missing in 25.36%1. A run can have more than one.

A better model did not make it more dependable

Claude Opus 5.5 topped the single-attempt scores at 67.16%1. It passed 241 of 507 tasks on all 20 attempts. Claude Opus 5, which scored lower, also passed 241. Half a point of headline accuracy bought, in Microsoft's words, "no additional dependability at all."

The spread is wider elsewhere. Kimi-K3 solved 476 tasks at least once, the broadest of any model. It got only 68 right on every attempt1.

Read that as a buyer. A demo shows one run. Your month-end close is thousands of them.

The failures are mostly boring

Microsoft tagged every failed run with one cause and averaged across models. Tool usage was 79.9% of failures. Wrong state updates were 10.3%1. Agents usually got far enough to try the work, then did not recover from a tool error, a failed precondition or an empty lookup.

That is good news. Boring failures have boring fixes. Microsoft lists four: check the end state before you commit, classify errors so retries hit only the recoverable ones, cut the tools down to what the workflow needs, and require human approval on changes you cannot cheaply reverse. They also say plainly they have not measured how much any of these helps1.

What to check this week

1. What your approver actually sees. Open a held action. Is it the agent's sentence, "invoice matched and ready to pay"? Or the change itself: this record, this field, from this value to that one? Only the second can be judged.

2. One case, twenty times. Pick a real case your agent handles. Run it twenty times against a test copy. Count how many end in the same correct state. The benchmark's whole point is that one pass proves almost nothing1.

3. What happens after a failed lookup. Most failures sat here. Find out what your agent does when a search comes back empty. Retry, try a listed alternative, or hold. Anything else should be refused.

4. Extra changes. 43.30% of clean-looking failures touched something they should not have1. Compare what the agent was asked to change with what it changed. Anything outside the list is a finding.

Evaluating agents in production covers building your own test set. Retries, idempotency and side effects covers the failed-call path.

What changes

For a year the question was "can the model do this?" ThinkingBox shows that is the wrong question. Most models can, once. Your question is whether it does the same thing on run two thousand. And whether you can see it when it doesn't.

That is what a control plane is for. In Gatehouse, each action is checked against a signed rules file before it runs. Anything outside it is refused, or held for one named approver. Every decision lands in the record.

Start with check 1. It takes ten minutes and tells you what your approver is really approving. Want a second pair of eyes on one live process? Book a teardown.

Sources

  1. [1]Microsoft and Hugging Face, The Agent Said It Was Done. The Database Disagreed. (3 October 2026)huggingface.co In text
  2. [2]One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows, arXiv:2608.19741arxiv.org
  3. [3]microsoft/thinkingbox-data, benchmark tasks and checks (GitHub)github.com

Want the next one? News with a take, three times a week. Follow by RSS

Read next

[ your next step ]

Bring us the queue nobody wants.

One process, studied in writing. You keep the document, whatever it says.

support@surehand.io