surehand
All articlesMeta

The stack for governed agents: six layers between a model and your ledger

Reference6 min readSurehand

A vendor tells you their agent handles accounts payable. Six pieces of software are involved. This page names them, says what each does, and marks the two that decide whether your controller sleeps.

In one sentence: between the model and your ledger sit orchestration, tools, control and record, and only the control layer must see every action.

The six layers

Read from the bottom up. The model is furthest from your business. The system of record is your business.

LayerWhat it doesWho usually sells itSwappable?
1. ModelReads text and produces text: a decision, a draft, a fieldOpenAI, Anthropic, Google, open-weight vendorsYes, and plan to
2. OrchestrationRuns the loop: prompt, call a tool, read the result, pick the next stepFrameworks, workflow tools, vendor platformsYes
3. ToolsWhat the agent can call: read an invoice, look up a PO, post a paymentConnectors, APIs, MCP servers, RPA botsYes
4. ControlChecks each action against rules before it runs. Holds the hard onesRarely sold on its own. Often a paragraph in a promptShould not be
5. RecordWrites down what happened, in a form an outsider can verifyLogging tools, observability platforms, or nothingShould not be
6. System of recordYour ERP, claims platform, ticketing system, vendor masterAlready yoursNo

Anthropic draws a useful line. Workflows are "systems where LLMs and tools are orchestrated through predefined code paths". Agents are "systems where LLMs dynamically direct their own processes and tool usage"1. Both live in layer 2. Both need layers 4 and 5 once layer 3 can write.

Layer by layer, in your nouns

1. The model. It reads the invoice. It says the bank details differ from the file. It reads and classifies well. It has never seen your delegation of authority. It will give a confident wrong answer. Treat it as a very fast junior who skipped the policy manual.

2. Orchestration. The loop. It feeds the model the invoice, runs the PO lookup, reads the result, and decides whether to schedule the payment. Most agent products live here. When a demo impresses you, you are watching layer 2.

3. Tools. Each tool is a capability. Reading a PO is harmless. Posting a payment is not. OWASP calls the failure excessive agency. Its root causes are "excessive functionality; excessive permissions; excessive autonomy"2. OWASP's own example: a developer needs read access to documents, and the extension they pick "also includes the ability to modify and delete documents"2. Your tool list is your blast radius.

4. Control. It stands between the model deciding to pay and the payment posting. It asks three questions of every action. Is this allowed at all? Is it allowed now, with these facts? Does a person look first? Security architects know this pattern. In NIST's zero trust model, access "is granted through a policy decision point (PDP) and corresponding policy enforcement point (PEP)"3. For agents, read "access" as "this action".

Most agent products put layer 4 in the prompt. That is a request. The model can ignore it. A well-crafted invoice can talk it out of it.

5. Record. What the run leaves behind. Not the application log, which is for engineers. A record your reviewer can read. Which rules were in force. What the agent saw. What it decided. Who approved. What changed. The AI Act asks high-risk systems to "technically allow for the automatic recording of events (logs) over the lifetime of the system"4. Your auditor will ask for more: proof nobody edited it since.

6. Systems of record. Your ERP, your claims platform, your ticketing tool. Nothing here is new. What is new: software with no human behind it now holds a write token. Every credential you hand to layer 3 is a credential to layer 6.

The pattern it borrows from

Kubernetes splits a cluster into "a control plane and one or more worker nodes". The control plane components "manage the overall state of the cluster"5. The nodes run the work. The thing that decides what may run is separate from the thing running it. It outlives any single job.

Agent stacks are rebuilding this split under pressure. Layer 2 is the worker. Layer 4 is the control plane. In Kubernetes nobody argues about whether the control plane should exist.

Where the risk lives

Layers 1 to 3 fail loudly. A bad output, a broken loop, a tool timeout. You see it in week one and fix it.

Layers 4 and 5 fail quietly. An agent with no control layer handles thousands of invoices well. Then it pays the one with the changed bank details. An agent with no record works fine until your auditor asks why a claim closed in March. Nobody can say which rules were in force.

So a demo tells you almost nothing. The demo is layers 1 to 3 on a clean example. Your deployment is layers 4 to 6 on your worst week.

What to check

When someone shows you an agent, ask which layers they are selling.

  1. /01

    Which layer does your product occupy? If they say all of them, ask to see layer 4 apart from layer 2.

  2. /02

    Where do the rules live? In a prompt, or in a component the model cannot rewrite?

  3. /03

    Who can approve a held action? Anyone with the link, or one named person?

  4. /04

    Show me the record of one run. Show me how I would prove nobody edited it.

  5. /05

    Which of my credentials does layer 3 hold, and with what scope?

  6. /06

    If I swap the model next year, which layers change?

Where it is going

Our view: layers 1 to 3 are becoming commodities. Every framework can call a tool. In two years nobody pays for orchestration alone. It comes with the model vendor or the workflow tool you already own.

Layers 4 and 5 go the other way. Regulators are writing them into law. Auditors are learning to ask. They encode your delegation of authority, not the technology. Expect control to become its own category with its own buyers, the way identity did.

Gatehouse fit

Gatehouse is layers 4 and 5 as one product. Signed, versioned rules. A check on every action before it runs. A named approver for held cases. A record chained by SHA-256 and exported in open formats. It is built to sit under whichever layer 2 you use, in front of your systems of record. Surehand deploys and runs it. See the Gatehouse page for the mechanism, and Compare for when a framework or workflow tool is the better answer.

At a glance

CategoryMeta (map)
LayersModel, orchestration, tools, control, record, system of record
Borrowed fromControl plane and worker nodes (Kubernetes). PDP and PEP (NIST zero trust)
Key sourcesAnthropic on workflows vs agents. OWASP LLM06. NIST SP 800-207. AI Act Art. 12
Typical ownerLayers 1-3: IT or the vendor. Layers 4-5: risk, internal audit, the process owner. Layer 6: already owned
The two to insist onControl and record

Sources

  1. [1]Building effective agents, Anthropic, December 2024anthropic.com In text
  2. [2]LLM06:2025 Excessive Agency, OWASP Top 10 for LLM Applicationsgenai.owasp.org In text
  3. [3]NIST SP 800-207, Zero Trust Architecture, August 2020nvlpubs.nist.gov In text
  4. [4]Regulation (EU) 2024/1689 (AI Act), Article 12: Record-keepingartificialintelligenceact.eu In text
  5. [5]Kubernetes Components, kubernetes.iokubernetes.io In text

Read next

[ your next step ]

Bring us the queue nobody wants.

One process, studied in writing. You keep the document, whatever it says.

support@surehand.io