
Shadow mode tests your AI agent on real work without letting it touch anything. The agent gets every live case. It decides what it would do and writes that down. Your people do the job as before. At the end you compare the two columns. Where they agree, you have evidence. Where they differ, you have a list.
In one sentence: the agent works your live queue with its hands tied, and you grade it against your own people.
Where it comes from
Machine learning teams have long run new models "in shadow". Amazon's SageMaker documents the pattern. The service "routes a copy of the inference requests" to the new variant "in real time within the same endpoint"1. "Only the responses of the production variant are returned to the calling application"1. The point is to "catch potential configuration errors and performance issues before they impact end users"1.
The idea moved to agents once agents started acting. In July 2026 Microsoft shipped Shadow Mode for its Dynamics 365 Case Management Agent. It lets you "evaluate AI on live production cases, without taking any action"2. The agent "observes, predicts, recommends, and simulates exactly what it would do if it were live"2. Microsoft lists what it will never do: update a case record, send a customer communication, change a case status, interrupt a workflow2.
Two things change when the shadowed thing is an agent. The output is an action, not a score. So you ask whether it would have paid this invoice, not whether the probability was close. And the human column is a person's decision. Slower, dearer, and not always right.
How it works
Take a claims triage agent at a third-party administrator. The numbers are illustrative.
Set-up. Connect the agent to the intake queue, read only. Replace its action tools (route to adjuster, request documents, close as duplicate) with one tool: write the proposed action and the reasons to a table. Your adjusters keep working the queue as last month.
Running. For four weeks, every claim gets two entries. The adjuster's actual action, from the claims platform. The agent's proposal, from the shadow table, with its evidence and the rule it thinks applies. The agent never sees the adjuster's choice first.
Reading the results. At the end you have something like this.
| Outcome | Count | What it tells you |
|---|---|---|
| Agent and adjuster agree | 1,610 | Your evidence for go-live |
| Agent would have held, adjuster acted | 190 | Your hold rate at these thresholds. Were the holds sensible? |
| Agent acted, adjuster held or escalated | 45 | The dangerous column. Read every one |
| Different actions | 155 | Split into agent wrong, adjuster wrong, both defensible |
The third row decides whether you go live. The fourth decides where the thresholds go. The second shows your approver's week.
Exit. Shadow mode ends when the criteria you wrote at the start are met. Not when someone gets impatient. A typical set: agreement above a chosen rate on cases the agent would act on alone. Zero cases in the dangerous column in the last two weeks. A hold rate your named approver can absorb.
What it is good at, and what it is not
Good at. Replacing the demo with evidence. A demo runs on ten clean examples. Shadow mode runs on your 2,000 real ones, Friday afternoons included. It finds your thresholds: you see what a 250.00 limit does to hold volume before anyone waits on an approval. It sizes the approver's workload. And every disagreement you review becomes a labelled example for later evaluation.
Not good at. Measuring correctness. It measures agreement with your people, and your people make mistakes. If adjusters wrongly close 3% of claims, an agent that gets those right shows up as 3% disagreement. Every disagreement needs a third opinion before it counts against the agent.
What it does not test. The write path. The agent never calls the real close-claim tool. So you learn nothing about timeouts, errors, or a write that lands twice. Your first live week will find those. Nor does it show how the queue behaves once the human column is gone.
A subtle trap. If adjusters see the agent's proposal before they decide, you stop measuring the agent. You measure how readily people agree with a suggestion. Keep the columns blind until the comparison.
What to check
Before you agree to a shadow period, get these in writing.
- /01
Is write access removed, or is the agent told not to act? Only the first is shadow mode.
- /02
What are the exit criteria? Who signs that they were met?
- /03
Who reviews disagreements? Are they independent of the people being compared?
- /04
Is the human column blind to the agent's proposal?
- /05
Are disagreements kept as labelled examples?
- /06
What happens to the write path after? A period with holds on every action, or straight to autonomy?
If you are a deployer of a high-risk system under the EU AI Act, shadow mode also feeds Article 26. Deployers must "monitor the operation of the high-risk AI system on the basis of the instructions for use"3. If they have reason to think it presents a risk, they inform the provider and "suspend the use of that system"3. A four-week shadow report is your first monitoring record. The disagreements table is what you attach.
Where it is going
Shadow mode is becoming a product feature rather than a project phase. Microsoft's July 2026 release is one. Our view: others will follow, because buyers keep asking to see the agent on live cases. The feature will matter less than the discipline around it. Blind comparison. Written exit criteria. Independent review of disagreements. A vendor can ship the toggle. Only you can decide what ready means.
Gatehouse fit
A Surehand engagement starts with a teardown of one queue before anything is built. It measures volume, exception rate and expected holds. That is the scoping a shadow period needs, and its thresholds are the ones the shadow period tests. Gatehouse keeps the disagreements you review as labelled examples, so your shadow table becomes the baseline for continuous evaluation once the agent is live. Read why AI projects fail for what happens when this step is skipped. Book a teardown if you want the table for one of your queues.
At a glance
| Category | Operations |
|---|---|
| Also called | Shadow deployment, shadow testing, parallel run |
| Borrowed from | ML model serving (SageMaker shadow variants). Parallel runs in system migrations |
| Key docs | SageMaker shadow tests. Dynamics 365 Shadow Mode (July 2026). AI Act Art. 26(5) |
| Typical owner | The process owner runs it. Risk reviews disagreements. The future approver reads the hold column |
| The one test | Is write access removed, or merely discouraged? |
Sources
- [1]Shadow tests, Amazon SageMaker AI Developer Guidedocs.aws.amazon.com In text
- [2]Trust before you automate: introducing Shadow Mode in Case Management Agent, Microsoft Dynamics 365 blog, 9 July 2026microsoft.com In text
- [3]Regulation (EU) 2024/1689 (AI Act), Article 26: Obligations of deployers of high-risk AI systemsartificialintelligenceact.eu In text
Read next

Why do AI projects fail?
An MIT NANDA report found most generative AI pilots showed no measurable return. The model was rarely the problem. The work around it was.
How to monitor AI agents in production
Uptime tells you it's running. Not that it's right. Four signals to watch: quality drift, hold rate, spend per run, latency.

Agent autonomy levels: where the human hold belongs
Autonomy isn't a setting for the whole agent. It's a choice you make per action. Put the hold where a mistake stops being cheap to undo.