
A hold sends a case to a person. It is the safety net. It is also someone's afternoon. Set the hold rate too high and your approver drowns. Too low and the agent decides things it should not. The maths that sits between those two is simple. Most teams never do it.
In one sentence: your approver's workload is volume times hold rate times minutes per hold, and you should know that number before go-live.
Where it comes from
Operations teams have always sized queues this way: items in, handling time, people available. The agent changes only what goes in the queue. Instead of every item, your people see the held ones.
The regulation adds a reason to care. If the AI Act applies, deployers must give oversight to people with "the necessary competence, training and authority, as well as the necessary support"2. It also wants overseers to stay aware of the risk of "automatically relying or over-relying on the output"3. An approver with no time is not supported. They are a rubber stamp.
How it works
Three inputs.
Volume. Items the agent handles per month. Invoices, claims, tickets.
Hold rate. The share of items that stop for a person. Your shadow period measures it. Before that, estimate it from today's exception rate.
Minutes per hold. How long your approver spends on one case. This depends on what arrives with the case.
Now a worked example. All numbers are illustrative.
| Input | AP example |
|---|---|
| Invoices per month | 3,000 |
| Hold rate | 10% |
| Holds per month | 300 |
| Working days per month | 21 |
| Holds per day | About 14 |
| Minutes per hold, evidence attached | 4 |
| Minutes per hold, approver must dig | 15 |
| Approver time per day, evidence attached | About 1 hour |
| Approver time per day, must dig | About 3.5 hours |
Same hold rate. The difference between an hour and half a day is the evidence. A hold that arrives with the invoice, the PO, the receipt, the supplier history and the rule that fired is a four-minute decision. A hold that says "please review INV 7731" sends the approver into three systems.
Now change the hold rate.
| Hold rate | Holds per day | Time per day (4 min each) | Time per day (15 min each) |
|---|---|---|---|
| 2% | About 3 | About 11 minutes | About 43 minutes |
| 5% | About 7 | About 29 minutes | About 1.8 hours |
| 10% | About 14 | About 1 hour | About 3.6 hours |
| 20% | About 29 | About 1.9 hours | About 7.1 hours |
At 20% with poor evidence, your approver has no other job. That is the point where holds stop being a control. Every case gets approved because there is no time to do anything else.
Reading the hold rate
The raw rate is less useful than what happens to the holds.
Approved almost always. If one trigger produces holds your approver approves nearly every time, the trigger is too tight. Loosen it, with a second signature, and the load falls.
Declined often. Good. The hold is catching real problems. Keep it.
Overturned in both directions. The rule is vague. Rewrite the trigger as a fact.
Growing month on month. Something changed: new suppliers, a new model, a new product line. Look before you add a second approver.
What it is good at, and what it is not
Good at. Turning "human in the loop" into a staffing number. Showing the business case honestly: the agent does 90% of the work, and one person spends an hour a day on the rest.
Not good at. Telling you whether the holds are the right ones. A low hold rate with bad triggers is cheap and dangerous. Pair this maths with the shadow table: are the held cases the ones a person should see?
The overload risk. Parasuraman and Manzey's review found that automation complacency "occurs under conditions of multiple-task load", and that automation bias "cannot be prevented by training or instructions"1. Workload is the lever you control. Size the queue so each hold gets real attention.
What to check
- /01
What is the expected hold rate, and where does the number come from?
- /02
How many minutes per hold, with the evidence as delivered? Time five real ones.
- /03
Who is the approver, and what else is in their day?
- /04
What is the plan when holds spike at month-end? Cover, extended time limit, or a paused queue?
- /05
Which triggers produce holds that are almost always approved?
- /06
Who reviews the hold rate each month?
Where it is going
Our view: approver load becomes a service level. Buyers will ask vendors for an expected hold rate and minutes per hold, the way they ask for uptime. Vendors that gather evidence into the hold will win on this number, because it is where the saving actually shows up.
Gatehouse fit
In Gatehouse a held case goes to one named approver with the evidence gathered before they are asked. That is the four-minute column, not the fifteen. Every decision is saved with the approver's name as a labelled example, so you can see which triggers are always approved and change the rule. A Surehand teardown estimates volume, exception rate and expected holds for one process before anything is built. Book a teardown to get your numbers.
At a glance
| Category | Cost |
|---|---|
| Formula | Volume x hold rate x minutes per hold |
| Biggest lever | Evidence delivered with the hold |
| Warning sign | A trigger whose holds are almost always approved |
| Key sources | AI Act Art. 14 and 26(2). Parasuraman and Manzey, 2010 |
| Typical owner | The process owner sizes it. The approver reports it. Risk reviews it monthly |
| The one test | How many minutes did your approver spend on holds yesterday? |
Sources
- [1]Parasuraman and Manzey, Complacency and bias in human use of automation, Human Factors, 2010 (PubMed abstract)pubmed.ncbi.nlm.nih.gov In text
- [2]Regulation (EU) 2024/1689 (AI Act), Article 26: Obligations of deployers of high-risk AI systemsartificialintelligenceact.eu In text
- [3]Regulation (EU) 2024/1689 (AI Act), Article 14: Human oversightartificialintelligenceact.eu In text
Read next

How a hold works: triggers, named approvers, cover and time limits
Human in the loop only works if each hold has four parts: a trigger, one named approver, a named cover, and a time limit. What each part does, and what happens when one is missing.

How much does an AI agent cost?
Three numbers. What it costs to build, what each run costs, and what it can commit in your name. Most quotes only give you the first.

Confidence thresholds for AI agents: what the score means, and why it drifts
'Hold anything below 80% confidence' only works if 80% means 80%. What calibration is, why language models often are not calibrated, and how to set a threshold you can defend.