surehand
Articles / Reference

Reference

One question, answered in full, for the people who have to sign off on an agent.

posts
27
updated
feed
rss
Vendors

LangGraph vs OpenAI Agents SDK vs CrewAI vs AutoGen: who gives you an approval step

Reference5 min read

Four popular agent frameworks, read from their own docs for one question: can a person approve an action before it runs, and what happens to the run while they decide?

Operations

Confidence thresholds for AI agents: what the score means, and why it drifts

Reference5 min read

'Hold anything below 80% confidence' only works if 80% means 80%. What calibration is, why language models often are not calibrated, and how to set a threshold you can defend.

Regulation

Data residency for AI agents: where your prompts, documents and records actually live

Reference5 min read

An agent sends your documents to a model, keeps state between steps, and writes a record. Each can sit in a different place. What the main providers' docs say, and what to ask.

Regulation

EU AI Act obligations for deployers of AI agents, after the 2026 Omnibus

Reference6 min read

If you use an AI agent rather than build one, you are a deployer. What Article 26 asks of you, which agents count as high-risk, and the dates the 2026 Omnibus moved.

Operations

Evaluating agents in production: a golden set for back-office work

Reference5 min read

How to know your agent still works after a model update, a new rule or a new supplier: a golden set of your own labelled cases, replayed before every change and sampled every week.

Cost

Frontier model vendors as suppliers: what to read in OpenAI's and Anthropic's terms

Reference5 min read

Your agent runs on a model you rent. Treat the model vendor like any critical supplier: price changes, retirement notice, data terms and exit. What their own terms and docs say.

Cost

Hold-rate maths: what one hold in ten costs your approver

Reference5 min read

The hold rate decides whether human in the loop is a control or a bottleneck. How to work out your approver's day from volume, hold rate and minutes per case, with a worked example.

Meta

How to read the Governed Agents Reference

Reference5 min read

What this reference is, what it is not, how the pages are built, and how to tell our opinion from the sourced facts.

Controls

How a hold works: triggers, named approvers, cover and time limits

Reference5 min read

Human in the loop only works if each hold has four parts: a trigger, one named approver, a named cover, and a time limit. What each part does, and what happens when one is missing.

Risks

Indirect prompt injection through documents: when the invoice gives the orders

Reference5 min read

Indirect prompt injection hides instructions in the documents your agent reads: invoices, claims, tickets, emails. What it is, why filters do not end it, and where the defence has to sit.

Regulation

ISO/IEC 42001 vs the NIST AI RMF, in plain English

Reference6 min read

One is a certifiable management-system standard. The other is a voluntary framework. What each asks of a company running AI agents, and what a certificate proves.

Controls

Kill switch, pause and rollback: what 'stop the agent' has to stop

Reference5 min read

A kill switch that stops new runs but not the one mid-payment is half a switch. What stop has to cover: in-flight actions, queued actions, scheduled runs and the credentials.

The product

The controls in these notes are what Gatehouse, Surehand’s control plane, enforces at runtime.

See Gatehouse
[ your next step ]

Bring us the queue nobody wants.

One process, studied in writing. You keep the document, whatever it says.

support@surehand.io