surehand
All articlesOperations

Evaluating agents in production: a golden set for back-office work

Reference5 min readSurehand

Your agent passed its tests in March. Since then the vendor swapped the model, someone edited the rules, and three new suppliers started sending invoices in a new layout. Does it still work? A golden set answers that. It is a fixed list of your own cases with the right answer written down. You replay it before every change, and it tells you what broke before your approver finds out.

In one sentence: keep a list of your real cases and their right answers, and never ship a change that makes the agent worse on them.

Where it comes from

Software teams call it regression testing. Machine learning teams call it an evaluation set. Anthropic's engineering team splits agent evals into two kinds. Capability evals ask whether the agent can do something new. Regression evals ask whether it still does what it used to. Tasks "that once measured 'Can we do this at all?' then measure 'Can we still do this reliably?'"1.

The same guide warns what happens without them. Teams get stuck "catching issues only in production, where fixing one failure creates others"1. It also sets a low bar to start: "20-50 simple tasks drawn from real failures is a great start"1.

Monitoring live work is older still. AWS describes Model Monitor as a way to watch models in production and detect drift in data and model quality2. NIST's AI RMF puts it under Measure: track the system's performance over time3.

How it works

A golden set for back-office work has four parts.

1. Cases. Real items from your queue, with personal data handled as your policy says. Invoices, claims, tickets. Pick them on purpose. The clean majority, yes, but also the hard ones: partial deliveries, credit notes, duplicate submissions, the supplier who always sends two PDFs.

2. Right answers. For each case, the action a competent person would take and why. Match and schedule. Hold for approval. Refuse. The best source is decisions your approver already made. Every hold they decided is a labelled example with a name on it.

3. Graders. How you score the agent's answer. Grade the outcome, not the wording. Anthropic's example: a flight-booking agent "might say 'Your flight has been booked'", but the outcome is "whether a reservation exists" in the database1. For AP, the question is not what the agent said. It is what it proposed to write to the ERP.

4. A rule for shipping. Before a model, prompt or rule change goes live, replay the set. Decide in advance what counts as a failure. For example: any case that should hold but would proceed blocks the release.

Here is an illustrative set for an AP agent.

SliceCasesRight answerWhy it is in the set
Clean three-way match20ProceedThe bulk of the queue
Over the autopay limit8Hold for AP leadTests the threshold
Bank details first seen recently6HoldTests the most expensive failure
Duplicate invoice, new PDF6Refuse, flag duplicateTests matching
Injected instruction in the PDF4Hold or refuseTests the rules, not the model
Credit note, partial delivery6Proceed with the right amountTests the edge cases people get wrong

Fifty cases. A person can read every result in an hour.

Production is not the golden set

A golden set is fixed. Your live queue is not. New suppliers, new claim types and seasonal peaks change the work under the agent. The golden set cannot see that.

So add a second loop: sample live cases every week and have a person grade them. Twenty is enough to spot a shift. When a sampled case is wrong in a new way, add it to the golden set. That is how the set grows from real failures, as Anthropic suggests.

Watch the numbers that move first. The hold rate. The share of holds your approver overturns. Cases the agent could not classify. A jump in any of them means the work changed, the agent changed, or both.

What it is good at, and what it is not

Good at. Making change safe. The vendor's model update, your rule tweak and a new prompt all get the same test. It also gives your auditor evidence that someone checks the agent, with dates.

Not good at. Proving the agent is right on cases nobody labelled. A golden set measures fifty cases, not fifty thousand. It is a smoke alarm, not a guarantee.

Easy to get wrong. Grading with another model and never checking the grader. Model graders are useful. Anthropic's teams pair them with periodic human calibration1. Read a sample of the grades yourself.

What to check

  1. /01

    Is there a golden set, and who owns it? How many cases, from where?

  2. /02

    Is it replayed before every model, prompt and rule change? Show me the last run.

  3. /03

    What result blocks a release?

  4. /04

    Are graders checking outcomes or wording?

  5. /05

    How are live cases sampled and graded each week?

  6. /06

    When did a new failure last get added to the set?

Where it is going

Our view: evaluation moves from the vendor's lab to the buyer's contract. Today you trust the vendor's benchmark. Soon you will ask for the replay of your own golden set before each model change, the way you ask for a change log. The vendors ready for that will keep customers through model churn.

Gatehouse fit

Gatehouse saves each approver decision as a labelled example. Before a model, prompt or policy change ships, it is replayed against those examples. That is a golden set built from your own approver's work, growing every week. The results sit next to the record of the runs they came from. The Gatehouse page shows where continuous evaluation fits.

At a glance

CategoryOperations
Also calledRegression evals, eval set, test set, golden dataset
Borrowed fromSoftware regression testing. ML evaluation
Key sourcesAnthropic, Demystifying evals for AI agents. NIST AI RMF Measure
Typical ownerThe process owner owns the right answers. IT or the vendor runs the replay
The one testShow me the last replay, and what would have blocked the release

Sources

  1. [1]Demystifying evals for AI agents, Anthropic Engineering, 9 January 2026anthropic.com In text
  2. [2]Data and model quality monitoring with Amazon SageMaker Model Monitor, AWS docsdocs.aws.amazon.com In text
  3. [3]NIST AI 100-1, AI Risk Management Framework 1.0, January 2023 (PDF)nvlpubs.nist.gov In text

Read next

[ your next step ]

Bring us the queue nobody wants.

One process, studied in writing. You keep the document, whatever it says.

support@surehand.io