surehand
All articlesGovernance

Agent autonomy levels: where the human hold belongs

Reference4 min readSurehand

"How autonomous should our agent be?" That question has no good answer. It uses the wrong unit. An agent isn't autonomous. Its actions are. Your claims agent can read a thousand documents a day and ask nobody. It should still never deny a claim on its own.

The levels, from the person's side

Most autonomy scales describe the machine. The useful ones describe you. A 2025 paper by Feng, McDonald and Zhang defines five levels by the user's role: operator, collaborator, consultant, approver, observer. Autonomy, it argues, should be "a deliberate design decision, separate from its capability and operational environment"1.

For back-office work, those five collapse into four:

LevelThe person...The agent...Typical back-office use
1. Assistdoes the worksuggests, drafts, looks things upDrafting a reply to a claimant
2. Checkreviews everything before it landsprepares complete workNew process, first weeks of a deployment
3. Approve by exceptiondecides only what the agent holdsacts when the rules allow, holds when they don'tInvoice matching once the hold rate is stable
4. Observereads the record afterwardsacts within scopeLow-value, easy-to-undo actions

Most pilots start at level 2. Many stall there. Reviewing everything works at a hundred items a week. At a thousand, your reviewer clicks approve without reading. Now you're at level 4, and nobody decided that. Our note on human in the loop covers that failure.

Set it per action

One agent, many actions, different levels. An illustrative claims triage agent at a third-party administrator:

ActionLevelWhy
Read claim documents and policy wording4. ObserveReading changes nothing
Classify claim type and route to a queue4. ObserveWrong queue costs a few hours and is easy to fix
Request missing documents from claimant3. Approve by exceptionCustomer-facing, but low stakes; hold if the claim is flagged
Recommend approval under a set amount3. Approve by exceptionHold anything near the limit or with a fraud signal
Deny a claim2. Check, alwaysHard to reverse, regulated, and a person answers for it
Change policy or claimant recordsNot allowedBelongs to another team

Illustrative. The amounts and flags come from your own authority limits.

Ask "is this agent level 3?" and you get an argument. Ask "is denying a claim level 3?" and you get an answer in a minute.

Where the hold belongs

One test. Put the hold where undoing a mistake stops being cheap.

Before that point, errors cost little. A misrouted claim goes back to the right queue. A draft email gets deleted. Let the agent move.

After that point, errors cost you. Three events mark the line in almost every back office:

  • Money leaves. A payment goes out. A refund is issued. A credit note posts. Getting it back means asking someone for it.
  • Someone outside hears. An email reaches a customer, supplier or regulator. You can't unsend it.
  • A record others rely on changes. The vendor master, the policy system, the general ledger. Other processes read it and act.

Hold just before each one. A hold at the start of the whole run is too early. It teaches your approver to wave runs through.

The EU AI Act asks the same of high-risk systems: people who can "decide, in any particular situation, not to use" the output, or "override or reverse" it2. Overriding only means something before the irreversible step. That's where the hold belongs.

How to raise autonomy without guessing

Levels should move. Move them on evidence.

  1. /01

    Start one level lower than you think. Most actions begin at level 2 or 3.

  2. /02

    Watch two numbers in the record. How often your approver overturns the agent. How often a proceeded action gets corrected later. No overturns on a class of decision for a month, and nothing reversed? That action can move up.

  3. /03

    Move one action at a time. Never the whole agent.

  4. /04

    Write down what would move it back. Say two corrections in a week drops it a level. Decide that now. Mid-incident is the wrong time.

  5. /05

    Re-check after model changes. A new model version can behave differently on the same inputs. Anthropic's own release notes for Opus 5.5 say the model "often suspects it is being evaluated"3. That makes lab results a weaker guide to live behaviour. Your record is the better guide. Our note on monitoring agents lists the signals.

Where Gatehouse fits

In Gatehouse, each action's level is one line in a signed rules file. The check runs before the action. A level-2 action can't quietly run at level 4. Holds go to one named approver. Every approve and decline becomes a labelled example. That's the evidence step 2 needs. Claims teams: the healthcare page shows this for triage.

Start here

List your agent's actions down the left of a page. For each one, ask: if this is wrong, what does it cost to undo? Mark the first action where the answer is "we'd have to ask someone for the money back" or "the customer already saw it". Your first hold goes there. Everything above it can probably run with less review than it gets now.

Sources

  1. [1]Feng, McDonald and Zhang, Levels of Autonomy for AI Agents, arXiv:2506.12469arxiv.org In text
  2. [2]Regulation (EU) 2024/1689 (AI Act), Article 14: Human oversightartificialintelligenceact.eu In text
  3. [3]Anthropic, Introducing Claude Opus 5.5 (22 September 2026)anthropic.com In text

Read next

[ your next step ]

Bring us the queue nobody wants.

One process, studied in writing. You keep the document, whatever it says.

support@surehand.io