Reference
One question, answered in full, for the people who have to sign off on an agent.
- posts
- 27
- updated
- feed
- rss

LangGraph vs OpenAI Agents SDK vs CrewAI vs AutoGen: who gives you an approval step
Four popular agent frameworks, read from their own docs for one question: can a person approve an action before it runs, and what happens to the run while they decide?

Confidence thresholds for AI agents: what the score means, and why it drifts
'Hold anything below 80% confidence' only works if 80% means 80%. What calibration is, why language models often are not calibrated, and how to set a threshold you can defend.

Data residency for AI agents: where your prompts, documents and records actually live
An agent sends your documents to a model, keeps state between steps, and writes a record. Each can sit in a different place. What the main providers' docs say, and what to ask.

EU AI Act obligations for deployers of AI agents, after the 2026 Omnibus
If you use an AI agent rather than build one, you are a deployer. What Article 26 asks of you, which agents count as high-risk, and the dates the 2026 Omnibus moved.

Evaluating agents in production: a golden set for back-office work
How to know your agent still works after a model update, a new rule or a new supplier: a golden set of your own labelled cases, replayed before every change and sampled every week.

Frontier model vendors as suppliers: what to read in OpenAI's and Anthropic's terms
Your agent runs on a model you rent. Treat the model vendor like any critical supplier: price changes, retirement notice, data terms and exit. What their own terms and docs say.

Hold-rate maths: what one hold in ten costs your approver
The hold rate decides whether human in the loop is a control or a bottleneck. How to work out your approver's day from volume, hold rate and minutes per case, with a worked example.

How to read the Governed Agents Reference
What this reference is, what it is not, how the pages are built, and how to tell our opinion from the sourced facts.

How a hold works: triggers, named approvers, cover and time limits
Human in the loop only works if each hold has four parts: a trigger, one named approver, a named cover, and a time limit. What each part does, and what happens when one is missing.

Indirect prompt injection through documents: when the invoice gives the orders
Indirect prompt injection hides instructions in the documents your agent reads: invoices, claims, tickets, emails. What it is, why filters do not end it, and where the defence has to sit.

ISO/IEC 42001 vs the NIST AI RMF, in plain English
One is a certifiable management-system standard. The other is a voluntary framework. What each asks of a company running AI agents, and what a certificate proves.

Kill switch, pause and rollback: what 'stop the agent' has to stop
A kill switch that stops new runs but not the one mid-payment is half a switch. What stop has to cover: in-flight actions, queued actions, scheduled runs and the credentials.
