surehand
All articlesControls

Segregation of duties for AI agents: when one agent initiates and approves

Reference5 min readSurehand

Segregation of duties is the oldest control in finance. The person who adds a supplier should not pay it. The person who pays should not reconcile the bank. Split the steps, and one person cannot commit an error and hide it. Now give all three steps to one AI agent. You have undone the split, and nobody signed off on it.

In one sentence: treat your agent as one person in the SoD matrix, and never let it hold both sides of a split.

Where it comes from

Accountants built it. Security engineers wrote it down. NIST's control catalogue carries it as AC-5, separation of duties. It "addresses the potential for abuse of authorized privileges and helps to reduce the risk of malevolent activity without collusion"1. It includes "dividing mission or business functions and support functions among different individuals or roles"1. And it warns that "separation of duty violations can span systems and application domains"1.

That last line is the agent problem. The agent does not break SoD inside one system. It breaks it across the ERP, the bank portal and the mailbox, because it holds credentials to all three.

Auditors expect alternatives when separation is hard. The PCAOB's standard for auditing internal control notes that a smaller company "might have fewer employees in the accounting function, limiting opportunities to segregate duties"2. It then says "the auditor should evaluate whether those alternative controls are effective"2. An agent is the same situation. The question is not whether duties are split. It is whether something else does the job.

How it works

Start with your SoD matrix for accounts payable. Put the agent in it as a row, like a person. The table is illustrative.

DutyTypical ownerAgent may hold it?
Add or change a supplier (vendor master)Master data teamNo
Enter a bill against a POAP clerkYes
Match bill, PO and receiptAP clerkYes
Approve payment over a thresholdAP lead or controllerNo. The agent proposes, a person approves
Release payment to the bankTreasuryNo
Reconcile the bankFinance, not APNo
Change the agent's rulesProcess owner with riskNo, and not the approver alone

Three principles fall out.

Give the agent one side of each split. It can prepare. It cannot approve what it prepared. OWASP's excessive agency entry names the failure: an agent that "fails to independently verify and approve high-impact actions"3. The fix is not a smarter agent. It is a different actor on the other side.

The rules are a duty. Someone writes what the agent may do. If the same person approves the agent's holds, they can loosen a rule on Monday and approve under it on Tuesday. Split writing the rules from approving the cases. Better still, need two names to loosen a rule.

The record is a duty. An agent that writes its own log can rewrite it. The component that records what happened should sit outside the agent's reach.

What it is good at, and what it is not

Good at. Giving your auditor a familiar frame. SoD is well understood. An agent mapped into the matrix is a new actor under an old control. Testing becomes possible: which duties does this identity hold, and is any pair incompatible?

Not good at. Catching a bad decision inside one duty. SoD stops the agent paying the supplier it created. It does not stop the agent matching the wrong invoice. That is what the hold and your approver are for.

Where it goes wrong. A shared service account. If the agent acts under a person's login, or one account for several bots, the matrix cannot tell who did what. NIST ties SoD to "account management activities in AC-2, access control mechanisms in AC-3, and identity management activities"1. No separate identity, no separation.

A trap in small teams. The person who built the agent often approves its holds and edits its rules. That is three duties in one person. In a small team it may be unavoidable. Then you need the alternative controls the PCAOB describes: a second reviewer on rule changes, a periodic look at approvals by someone else.

What to check

  1. /01

    Does the agent have its own identity in every system it touches?

  2. /02

    List the agent's duties across all systems. Are any two on your incompatible list?

  3. /03

    Can the agent approve, release or reconcile anything it prepared?

  4. /04

    Who can change the agent's rules? Is that the same person who approves its holds?

  5. /05

    Who can change the record? Is the agent excluded?

  6. /06

    If the AI Act applies, are the oversight people competent and authorised? Article 26(2) asks for exactly that4.

Where it is going

Auditors are starting to ask about agents in walkthroughs. Our view: within two years an agent identity will appear in the SoD matrix as routinely as a service account does today. Teams that gave each agent its own identity and a written list of duties will pass. Teams that ran agents on a senior accountant's login will spend the audit explaining.

Gatehouse fit

Gatehouse keeps the three sides apart. The agent acts under its own role, set in the signed rules, so its duties are listed and testable. Actions over a limit go to one named approver who is not the agent. Overriding a block takes two names. Changes to the rules only take effect once signed. The record is chained by SHA-256 outside the agent's reach. The composite invoice matching deployment shows the split in an AP run.

At a glance

CategoryControls
Also calledSeparation of duties, SoD, incompatible duties
Borrowed fromAccounting internal control. NIST SP 800-53 AC-5
Key standards or docsNIST AC-5, AC-6. PCAOB AS 2201. OWASP LLM06:2025
Typical ownerInternal audit or the controller owns the matrix. IT enforces access
The one testCan the agent approve, release or record anything it prepared?

Sources

  1. [1]NIST SP 800-53 Rev. 5, AC-5 Separation of Duties and AC-6 Least Privilege (PDF), September 2020nvlpubs.nist.gov In text
  2. [2]AS 2201: An Audit of Internal Control Over Financial Reporting, PCAOBpcaobus.org In text
  3. [3]LLM06:2025 Excessive Agency, OWASP Top 10 for LLM Applicationsgenai.owasp.org In text
  4. [4]Regulation (EU) 2024/1689 (AI Act), Article 26: Obligations of deployers of high-risk AI systemsartificialintelligenceact.eu In text

Read next

[ your next step ]

Bring us the queue nobody wants.

One process, studied in writing. You keep the document, whatever it says.

support@surehand.io