surehand
All articlesOperations

What is shadow mode for AI agents?

Reference5 min readSurehand

Shadow mode tests your AI agent on real work without letting it touch anything. The agent gets every live case. It decides what it would do and writes that down. Your people do the job as before. At the end you compare the two columns. Where they agree, you have evidence. Where they differ, you have a list.

In one sentence: the agent works your live queue with its hands tied, and you grade it against your own people.

Where it comes from

Machine learning teams have long run new models "in shadow". Amazon's SageMaker documents the pattern. The service "routes a copy of the inference requests" to the new variant "in real time within the same endpoint"1. "Only the responses of the production variant are returned to the calling application"1. The point is to "catch potential configuration errors and performance issues before they impact end users"1.

The idea moved to agents once agents started acting. In July 2026 Microsoft shipped Shadow Mode for its Dynamics 365 Case Management Agent. It lets you "evaluate AI on live production cases, without taking any action"2. The agent "observes, predicts, recommends, and simulates exactly what it would do if it were live"2. Microsoft lists what it will never do: update a case record, send a customer communication, change a case status, interrupt a workflow2.

Two things change when the shadowed thing is an agent. The output is an action, not a score. So you ask whether it would have paid this invoice, not whether the probability was close. And the human column is a person's decision. Slower, dearer, and not always right.

How it works

Take a claims triage agent at a third-party administrator. The numbers are illustrative.

Set-up. Connect the agent to the intake queue, read only. Replace its action tools (route to adjuster, request documents, close as duplicate) with one tool: write the proposed action and the reasons to a table. Your adjusters keep working the queue as last month.

Running. For four weeks, every claim gets two entries. The adjuster's actual action, from the claims platform. The agent's proposal, from the shadow table, with its evidence and the rule it thinks applies. The agent never sees the adjuster's choice first.

Reading the results. At the end you have something like this.

OutcomeCountWhat it tells you
Agent and adjuster agree1,610Your evidence for go-live
Agent would have held, adjuster acted190Your hold rate at these thresholds. Were the holds sensible?
Agent acted, adjuster held or escalated45The dangerous column. Read every one
Different actions155Split into agent wrong, adjuster wrong, both defensible

The third row decides whether you go live. The fourth decides where the thresholds go. The second shows your approver's week.

Exit. Shadow mode ends when the criteria you wrote at the start are met. Not when someone gets impatient. A typical set: agreement above a chosen rate on cases the agent would act on alone. Zero cases in the dangerous column in the last two weeks. A hold rate your named approver can absorb.

What it is good at, and what it is not

Good at. Replacing the demo with evidence. A demo runs on ten clean examples. Shadow mode runs on your 2,000 real ones, Friday afternoons included. It finds your thresholds: you see what a 250.00 limit does to hold volume before anyone waits on an approval. It sizes the approver's workload. And every disagreement you review becomes a labelled example for later evaluation.

Not good at. Measuring correctness. It measures agreement with your people, and your people make mistakes. If adjusters wrongly close 3% of claims, an agent that gets those right shows up as 3% disagreement. Every disagreement needs a third opinion before it counts against the agent.

What it does not test. The write path. The agent never calls the real close-claim tool. So you learn nothing about timeouts, errors, or a write that lands twice. Your first live week will find those. Nor does it show how the queue behaves once the human column is gone.

A subtle trap. If adjusters see the agent's proposal before they decide, you stop measuring the agent. You measure how readily people agree with a suggestion. Keep the columns blind until the comparison.

What to check

Before you agree to a shadow period, get these in writing.

  1. /01

    Is write access removed, or is the agent told not to act? Only the first is shadow mode.

  2. /02

    What are the exit criteria? Who signs that they were met?

  3. /03

    Who reviews disagreements? Are they independent of the people being compared?

  4. /04

    Is the human column blind to the agent's proposal?

  5. /05

    Are disagreements kept as labelled examples?

  6. /06

    What happens to the write path after? A period with holds on every action, or straight to autonomy?

If you are a deployer of a high-risk system under the EU AI Act, shadow mode also feeds Article 26. Deployers must "monitor the operation of the high-risk AI system on the basis of the instructions for use"3. If they have reason to think it presents a risk, they inform the provider and "suspend the use of that system"3. A four-week shadow report is your first monitoring record. The disagreements table is what you attach.

Where it is going

Shadow mode is becoming a product feature rather than a project phase. Microsoft's July 2026 release is one. Our view: others will follow, because buyers keep asking to see the agent on live cases. The feature will matter less than the discipline around it. Blind comparison. Written exit criteria. Independent review of disagreements. A vendor can ship the toggle. Only you can decide what ready means.

Gatehouse fit

A Surehand engagement starts with a teardown of one queue before anything is built. It measures volume, exception rate and expected holds. That is the scoping a shadow period needs, and its thresholds are the ones the shadow period tests. Gatehouse keeps the disagreements you review as labelled examples, so your shadow table becomes the baseline for continuous evaluation once the agent is live. Read why AI projects fail for what happens when this step is skipped. Book a teardown if you want the table for one of your queues.

At a glance

CategoryOperations
Also calledShadow deployment, shadow testing, parallel run
Borrowed fromML model serving (SageMaker shadow variants). Parallel runs in system migrations
Key docsSageMaker shadow tests. Dynamics 365 Shadow Mode (July 2026). AI Act Art. 26(5)
Typical ownerThe process owner runs it. Risk reviews disagreements. The future approver reads the hold column
The one testIs write access removed, or merely discouraged?

Sources

  1. [1]Shadow tests, Amazon SageMaker AI Developer Guidedocs.aws.amazon.com In text
  2. [2]Trust before you automate: introducing Shadow Mode in Case Management Agent, Microsoft Dynamics 365 blog, 9 July 2026microsoft.com In text
  3. [3]Regulation (EU) 2024/1689 (AI Act), Article 26: Obligations of deployers of high-risk AI systemsartificialintelligenceact.eu In text

Read next

[ your next step ]

Bring us the queue nobody wants.

One process, studied in writing. You keep the document, whatever it says.

support@surehand.io