
"How autonomous should our agent be?" That question has no good answer. It uses the wrong unit. An agent isn't autonomous. Its actions are. Your claims agent can read a thousand documents a day and ask nobody. It should still never deny a claim on its own.
The levels, from the person's side
Most autonomy scales describe the machine. The useful ones describe you. A 2025 paper by Feng, McDonald and Zhang defines five levels by the user's role: operator, collaborator, consultant, approver, observer. Autonomy, it argues, should be "a deliberate design decision, separate from its capability and operational environment"1.
For back-office work, those five collapse into four:
| Level | The person... | The agent... | Typical back-office use |
|---|---|---|---|
| 1. Assist | does the work | suggests, drafts, looks things up | Drafting a reply to a claimant |
| 2. Check | reviews everything before it lands | prepares complete work | New process, first weeks of a deployment |
| 3. Approve by exception | decides only what the agent holds | acts when the rules allow, holds when they don't | Invoice matching once the hold rate is stable |
| 4. Observe | reads the record afterwards | acts within scope | Low-value, easy-to-undo actions |
Most pilots start at level 2. Many stall there. Reviewing everything works at a hundred items a week. At a thousand, your reviewer clicks approve without reading. Now you're at level 4, and nobody decided that. Our note on human in the loop covers that failure.
Set it per action
One agent, many actions, different levels. An illustrative claims triage agent at a third-party administrator:
| Action | Level | Why |
|---|---|---|
| Read claim documents and policy wording | 4. Observe | Reading changes nothing |
| Classify claim type and route to a queue | 4. Observe | Wrong queue costs a few hours and is easy to fix |
| Request missing documents from claimant | 3. Approve by exception | Customer-facing, but low stakes; hold if the claim is flagged |
| Recommend approval under a set amount | 3. Approve by exception | Hold anything near the limit or with a fraud signal |
| Deny a claim | 2. Check, always | Hard to reverse, regulated, and a person answers for it |
| Change policy or claimant records | Not allowed | Belongs to another team |
Illustrative. The amounts and flags come from your own authority limits.
Ask "is this agent level 3?" and you get an argument. Ask "is denying a claim level 3?" and you get an answer in a minute.
Where the hold belongs
One test. Put the hold where undoing a mistake stops being cheap.
Before that point, errors cost little. A misrouted claim goes back to the right queue. A draft email gets deleted. Let the agent move.
After that point, errors cost you. Three events mark the line in almost every back office:
- Money leaves. A payment goes out. A refund is issued. A credit note posts. Getting it back means asking someone for it.
- Someone outside hears. An email reaches a customer, supplier or regulator. You can't unsend it.
- A record others rely on changes. The vendor master, the policy system, the general ledger. Other processes read it and act.
Hold just before each one. A hold at the start of the whole run is too early. It teaches your approver to wave runs through.
The EU AI Act asks the same of high-risk systems: people who can "decide, in any particular situation, not to use" the output, or "override or reverse" it2. Overriding only means something before the irreversible step. That's where the hold belongs.
How to raise autonomy without guessing
Levels should move. Move them on evidence.
- /01
Start one level lower than you think. Most actions begin at level 2 or 3.
- /02
Watch two numbers in the record. How often your approver overturns the agent. How often a proceeded action gets corrected later. No overturns on a class of decision for a month, and nothing reversed? That action can move up.
- /03
Move one action at a time. Never the whole agent.
- /04
Write down what would move it back. Say two corrections in a week drops it a level. Decide that now. Mid-incident is the wrong time.
- /05
Re-check after model changes. A new model version can behave differently on the same inputs. Anthropic's own release notes for Opus 5.5 say the model "often suspects it is being evaluated"3. That makes lab results a weaker guide to live behaviour. Your record is the better guide. Our note on monitoring agents lists the signals.
Where Gatehouse fits
In Gatehouse, each action's level is one line in a signed rules file. The check runs before the action. A level-2 action can't quietly run at level 4. Holds go to one named approver. Every approve and decline becomes a labelled example. That's the evidence step 2 needs. Claims teams: the healthcare page shows this for triage.
Start here
List your agent's actions down the left of a page. For each one, ask: if this is wrong, what does it cost to undo? Mark the first action where the answer is "we'd have to ask someone for the money back" or "the customer already saw it". Your first hold goes there. Everything above it can probably run with less review than it gets now.
Sources
- [1]Feng, McDonald and Zhang, Levels of Autonomy for AI Agents, arXiv:2506.12469arxiv.org In text
- [2]Regulation (EU) 2024/1689 (AI Act), Article 14: Human oversightartificialintelligenceact.eu In text
- [3]Anthropic, Introducing Claude Opus 5.5 (22 September 2026)anthropic.com In text
Read next

What is human in the loop AI?
It works when one person you name reviews only what the system is unsure of. You set where “unsure” starts.

What is an agent approval policy?
An approval policy decides which agent actions go ahead, which wait for a person, and which never happen. Most teams have a paragraph. You need a table.
How to monitor AI agents in production
Uptime tells you it's running. Not that it's right. Four signals to watch: quality drift, hold rate, spend per run, latency.