
Your agent passed its tests in March. Since then the vendor swapped the model, someone edited the rules, and three new suppliers started sending invoices in a new layout. Does it still work? A golden set answers that. It is a fixed list of your own cases with the right answer written down. You replay it before every change, and it tells you what broke before your approver finds out.
In one sentence: keep a list of your real cases and their right answers, and never ship a change that makes the agent worse on them.
Where it comes from
Software teams call it regression testing. Machine learning teams call it an evaluation set. Anthropic's engineering team splits agent evals into two kinds. Capability evals ask whether the agent can do something new. Regression evals ask whether it still does what it used to. Tasks "that once measured 'Can we do this at all?' then measure 'Can we still do this reliably?'"1.
The same guide warns what happens without them. Teams get stuck "catching issues only in production, where fixing one failure creates others"1. It also sets a low bar to start: "20-50 simple tasks drawn from real failures is a great start"1.
Monitoring live work is older still. AWS describes Model Monitor as a way to watch models in production and detect drift in data and model quality2. NIST's AI RMF puts it under Measure: track the system's performance over time3.
How it works
A golden set for back-office work has four parts.
1. Cases. Real items from your queue, with personal data handled as your policy says. Invoices, claims, tickets. Pick them on purpose. The clean majority, yes, but also the hard ones: partial deliveries, credit notes, duplicate submissions, the supplier who always sends two PDFs.
2. Right answers. For each case, the action a competent person would take and why. Match and schedule. Hold for approval. Refuse. The best source is decisions your approver already made. Every hold they decided is a labelled example with a name on it.
3. Graders. How you score the agent's answer. Grade the outcome, not the wording. Anthropic's example: a flight-booking agent "might say 'Your flight has been booked'", but the outcome is "whether a reservation exists" in the database1. For AP, the question is not what the agent said. It is what it proposed to write to the ERP.
4. A rule for shipping. Before a model, prompt or rule change goes live, replay the set. Decide in advance what counts as a failure. For example: any case that should hold but would proceed blocks the release.
Here is an illustrative set for an AP agent.
| Slice | Cases | Right answer | Why it is in the set |
|---|---|---|---|
| Clean three-way match | 20 | Proceed | The bulk of the queue |
| Over the autopay limit | 8 | Hold for AP lead | Tests the threshold |
| Bank details first seen recently | 6 | Hold | Tests the most expensive failure |
| Duplicate invoice, new PDF | 6 | Refuse, flag duplicate | Tests matching |
| Injected instruction in the PDF | 4 | Hold or refuse | Tests the rules, not the model |
| Credit note, partial delivery | 6 | Proceed with the right amount | Tests the edge cases people get wrong |
Fifty cases. A person can read every result in an hour.
Production is not the golden set
A golden set is fixed. Your live queue is not. New suppliers, new claim types and seasonal peaks change the work under the agent. The golden set cannot see that.
So add a second loop: sample live cases every week and have a person grade them. Twenty is enough to spot a shift. When a sampled case is wrong in a new way, add it to the golden set. That is how the set grows from real failures, as Anthropic suggests.
Watch the numbers that move first. The hold rate. The share of holds your approver overturns. Cases the agent could not classify. A jump in any of them means the work changed, the agent changed, or both.
What it is good at, and what it is not
Good at. Making change safe. The vendor's model update, your rule tweak and a new prompt all get the same test. It also gives your auditor evidence that someone checks the agent, with dates.
Not good at. Proving the agent is right on cases nobody labelled. A golden set measures fifty cases, not fifty thousand. It is a smoke alarm, not a guarantee.
Easy to get wrong. Grading with another model and never checking the grader. Model graders are useful. Anthropic's teams pair them with periodic human calibration1. Read a sample of the grades yourself.
What to check
- /01
Is there a golden set, and who owns it? How many cases, from where?
- /02
Is it replayed before every model, prompt and rule change? Show me the last run.
- /03
What result blocks a release?
- /04
Are graders checking outcomes or wording?
- /05
How are live cases sampled and graded each week?
- /06
When did a new failure last get added to the set?
Where it is going
Our view: evaluation moves from the vendor's lab to the buyer's contract. Today you trust the vendor's benchmark. Soon you will ask for the replay of your own golden set before each model change, the way you ask for a change log. The vendors ready for that will keep customers through model churn.
Gatehouse fit
Gatehouse saves each approver decision as a labelled example. Before a model, prompt or policy change ships, it is replayed against those examples. That is a golden set built from your own approver's work, growing every week. The results sit next to the record of the runs they came from. The Gatehouse page shows where continuous evaluation fits.
At a glance
| Category | Operations |
|---|---|
| Also called | Regression evals, eval set, test set, golden dataset |
| Borrowed from | Software regression testing. ML evaluation |
| Key sources | Anthropic, Demystifying evals for AI agents. NIST AI RMF Measure |
| Typical owner | The process owner owns the right answers. IT or the vendor runs the replay |
| The one test | Show me the last replay, and what would have blocked the release |
Sources
- [1]Demystifying evals for AI agents, Anthropic Engineering, 9 January 2026anthropic.com In text
- [2]Data and model quality monitoring with Amazon SageMaker Model Monitor, AWS docsdocs.aws.amazon.com In text
- [3]NIST AI 100-1, AI Risk Management Framework 1.0, January 2023 (PDF)nvlpubs.nist.gov In text
Read next
How to monitor AI agents in production
Uptime tells you it's running. Not that it's right. Four signals to watch: quality drift, hold rate, spend per run, latency.

What is shadow mode for AI agents?
Shadow mode runs the agent on live work while your people still do the job. The agent decides, nothing happens, and you compare. It is the cheapest go-live test there is.

Confidence thresholds for AI agents: what the score means, and why it drifts
'Hold anything below 80% confidence' only works if 80% means 80%. What calibration is, why language models often are not calibrated, and how to set a threshold you can defend.