Ask about scope, failure, cost and evidence, in that order, and ask for a single real run rather than a demonstration. Accuracy questions sound rigorous and tell you little, because every vendor's number was measured on data they chose. What separates a system you can operate from one you cannot is whether it can show you what it did.
Most evaluation calls follow the vendor's plan: a demonstration, a slide about the model and an accuracy figure. None of it is dishonest, but none of it shows what happens in your back office on an ordinary Monday. These twelve questions do.
Scope: what it is allowed to touch
- /01
Where are this agent's permissions written down, and can I read that document? If the answer is a conversation rather than a file, there is no scope, only habit.
- /02
What can it do that I have not explicitly approved? The honest answer is never nothing, and a vendor who says nothing has not thought about it.
- /03
How do I change scope after launch, and does that change leave a record? Scope you cannot narrow on Friday afternoon is not a control.
Failure: what happens when it is wrong
Every model is wrong sometimes. A vendor who will not say so is the wrong vendor.
- /01
What happens when your system is not sure? Listen for a threshold and a named approver. Listen for whether it stops at all.
- /02
Show me a run that was held, and who it went to.
- /03
How would I discover that quality had drifted, and how long would that take? If the answer relies on somebody noticing, the answer is months.
Cost: what it costs when nobody is watching
- /01
What is the per run cost, and what makes it go up? Retries, long documents and chatty prompts are the usual answers.
- /02
Can the system exceed a limit I set, and what stops it? A dashboard is not a limit. A limit stops the run.
- /03
What can it commit on my behalf, in money? This is separate from what it costs to run, and it is the number to cap first.
Evidence: what it can prove afterwards
- /01
Show me one complete run from a date I choose, rather than a prepared example. How the vendor responds tells you a lot.
- /02
What did that decision rest on, and can I read the source myself? Ask for the cited source itself.
- /03
If we part company, what do I keep and in what format? If you cannot export the trail without the vendor, you do not control it.
GatehouseThat is the sealed record of RUN-7F2K4, the simulated run on Gatehouse, Surehand's control plane. It reads step by step on the product site. Read the run on Gatehouse
How to read the answers
| You asked | Worth trusting | Worth worrying about |
|---|---|---|
| What is it allowed to touch? | A written manifest, per deployment | A description of good intentions |
| What if it is not sure? | A threshold and a named approver | It retries, or it flags for review |
| Who saw this decision? | A person, by role, in the record | The system logged it |
| What can it spend? | A hard limit, enforced in the path | We monitor usage closely |
| Show me one run | They open a real one, unprepared | They open the demo again |
| What do we keep? | The trail and policy history, in open formats | We can discuss that at renewal |
The right hand column is not proof of a bad vendor. It is proof that the question has not been answered yet, which is useful to know before signing rather than after.
The one question, if you only ask one
What happens when your system is not sure? Everything worth knowing follows from it. A vendor built for production answers with a behaviour and a person. Below a set confidence the run holds and goes to the named approver. A vendor who has built for demonstrations will answer with a description of how rarely that happens.
Ask us the same twelve. Our own twelve answers are on the trust page. Where we can only show the simulated run, we say so.
Keep reading

What should an AI audit trail contain?
Seven fields, and why logs are not an audit trail. What your buyer's risk function will ask for, and what most systems cannot produce.
What is AI governance?
Six controls that decide what an AI system may do, on whose authority, at what cost, and with what evidence. A practical guide rather than a policy template.