
Logs were built so engineers could find out why software broke. They were never built to prove what a business decided. For twenty years that didn't matter. Software didn't decide much. Now agents approve invoices, deny claims and email customers. Most of them leave behind nothing better than logs. That's the patch. The fix is a sealed record of every decision.
The question changed
When a web server fails, the engineer on call asks: what happened, in what order, and where did it go wrong? Logs answer that well. They're cheap. They're everywhere. The code writes them as it runs. NIST's guide to log management treats them as a security and operations tool. That's exactly what they are3.
When an agent pays the wrong supplier, someone else asks. Your controller. Your auditor. The customer. Sometimes a regulator. They don't ask what happened in the code. They ask:
- What did the agent decide?
- What was it allowed to do at the time?
- What did it look at?
- Who could have stopped it, and did they?
- Has anyone changed this account since?
Logs weren't designed for any of those questions. They can sometimes be made to answer them. By an engineer. With effort. If the right lines were written and nobody rotated them away. That's what a patch looks like. A tool built for one job, stretched to cover another. It works until the day it's tested.
What the patch costs
Relying on logs costs nothing in a normal week. The bill arrives the first time someone outside the team asks what happened.
Look at the largest example so far. OpenAI has spent the summer reviewing what its own agents did on the internet during training and evaluation. Its September update said it had notified dozens of third parties. Most cases so far were low severity. And "given the scale of the review required, and the need to verify each case, this work will take months to complete". It is working backwards "month by month"1. Australia's prime minister said publicly that OpenAI took too long to disclose an incident from June2.
This isn't a criticism of OpenAI's engineers. It happens to any team when the evidence is scattered across systems built for debugging. Each question becomes an investigation. Each investigation needs the few people who understand the systems. The answer arrives weeks later. The people asking can't check it for themselves.
Now scale it down to your business. Your approver asks why an invoice for 18,400 (an illustrative figure) was paid to a new account in March. With logs, someone searches three systems. Lines up timestamps. Guesses which prompt version was live. Writes a summary. The summary is their account of events. It isn't evidence.
Five ways logs fail when an agent acts
1. They're fragments. One decision leaves lines in the agent framework, the model provider, the ERP, the email system. Nothing ties them together except a timestamp and hope.
2. They record events. A log says a function was called with some arguments. It rarely says which rule allowed it. Or what threshold applied. Or what evidence was used. Or who was supposed to approve it.
3. They're written for insiders. A log line makes sense to the person who wrote the code. An auditor can't read it. And shouldn't have to.
4. They can be changed without a trace. Most log stores let an administrator edit, delete or drop entries. Rotation deletes them by design after a few weeks. Fine for debugging. Fatal for evidence. The person asking can't tell what's missing.
5. They stay with the vendor. If the agent runs on someone else's platform, the logs are theirs too. When the contract ends, so does your history.
None of this is new
Other fields hit this problem long before agents. They fixed it.
Brokers and dealers in the US must keep electronic records in a way that either preserves "a complete time-stamped audit trail" of every change and deletion, or in "non-rewriteable, non-erasable format"5. Regulators had worked it out. A firm can't be the only guard on records it might want to change.
The web's certificate system had a version of it. A certificate authority could issue a bad certificate and nobody outside would know. The answer, Certificate Transparency, was "publicly auditable, append-only" logs built on hash trees. The logs don't stop bad certificates. They make sure anyone can detect one4. Computer scientists had already described efficient tamper-evident logging in 20098.
The lesson from both. When a party could be tempted or pressured to rewrite history, you don't ask them to promise not to. You make rewriting visible.
Agents that act are exactly that case. The vendor, the operator and the client all have reasons to prefer a kinder version of events once something goes wrong.
What a sealed record is
A sealed record has four properties. We call it the SEAL test.
| Property | What it means | What logs usually do |
|---|---|---|
| Single decision | One record per run, complete, with everything needed to understand it | Many fragments per run, across systems |
| Entered at the time | Written when the run closes, not reconstructed later | Assembled after someone asks |
| Any edit shows | Each record carries a hash of the one before it, so any edit or deletion breaks the chain where anyone can see | Editable and rotated |
| Legible without us | Exports in open formats a reader can check with standard tools | Lives in the vendor's console |
What goes inside the record matters too. Our note on what an AI audit trail should contain lists seven fields: identity, scope, input, evidence, decision against the threshold, cost, and the person. The SEAL test is about the container. The seven fields are about the contents. You need both.
The hash chain is the part people ask about. It's simple. Each record includes a SHA-256 hash of the previous record. Change one byte in a record from March. Every hash after it stops matching. You don't have to trust the operator. You rerun the hashes yourself.
What sealing doesn't do
The honest limits matter. This idea gets oversold.
A seal doesn't prove the decision was right. It proves the record hasn't changed since it was sealed. A wrong decision, faithfully recorded, is still wrong. For that you need rules and a named approver in front of the action. That's what an approval policy is for.
A seal doesn't prove the record was complete when sealed. If the system never captured the evidence, sealing preserves the gap. So the record has to be written by the same layer that checks the action. Records collected from logs afterwards inherit the logs' gaps.
A seal isn't a certification. It's a property of the data. It says nothing about how the company is run. We say plainly on our trust page that Surehand doesn't yet hold SOC 2 or ISO 27001. A hash chain doesn't replace either.
It doesn't replace logs. Engineers still need them. Keep your logs for debugging. Just stop asking them to be evidence.
Why now
Three things make this urgent rather than tidy.
First, agents are acting outside the lab. The OpenAI disclosures show that even the best-funded teams find out what their agents did long after the fact1.
Second, rules are arriving. The EU AI Act requires high-risk systems to "technically allow for the automatic recording of events (logs) over the lifetime of the system"6. It requires deployers to keep those logs for a period suited to the purpose, "of at least six months"7. Where a use counts as high-risk under Annex III, those duties now start on 2 December 2027. Many back-office agents won't be high-risk under the Act at all. Your auditor won't care. The law says "logs". An auditor reads that as evidence.
Third, records compound. A sealed record from today is the evidence that lets you give an agent more autonomy next quarter. A log from today will be rotated away before anyone needs it.
What this looks like in practice
In Gatehouse, the record is written by the same layer that checks each action against the signed rules file. Every run is saved as one record. Each is chained to the one before it with SHA-256. Each exports in open formats. Your approver's decision is part of the record. It's kept as a labelled example for tuning the rules later. If you leave, the records leave with you.
You don't need Gatehouse to start. You need to find out whether you have a record or a patch.
A ten-minute test
Pick one agent that approves, pays, sends or changes records. Pick a run from last month at random. Let nobody prepare it. Then ask for:
- /01
What it decided, and the rule that allowed it.
- /02
What it read to decide.
- /03
Who could have stopped it.
- /04
Proof that nobody has changed any of that since.
Time it. More than ten minutes? Needs an engineer? Fourth answer is "you'll have to trust us"? You're running on a patch. Our note on how to monitor AI agents covers what to track once the records exist. Want help running the test on a live process? book a teardown.
Sources
- [1]OpenAI, The Hugging Face incident and other third-party impact from misaligned modelsopenai.com In text
- [2]BBC News, Rogue OpenAI agent 'infiltrated' Australian government website (24 September 2026)bbc.com In text
- [3]NIST SP 800-92, Guide to Computer Security Log Managementcsrc.nist.gov In text
- [4]RFC 6962, Certificate Transparencyrfc-editor.org In text
- [5]17 CFR 240.17a-4, Records to be preserved by certain exchange members, brokers and dealerslaw.cornell.edu In text
- [6]Regulation (EU) 2024/1689 (AI Act), Article 12: Record-keepingartificialintelligenceact.eu In text
- [7]Regulation (EU) 2024/1689 (AI Act), Article 26: Obligations of deployers of high-risk AI systemsartificialintelligenceact.eu In text
- [8]Crosby and Wallach, Efficient Data Structures for Tamper-Evident Logging, USENIX Security 2009usenix.org In text
Read next

What should an AI audit trail contain?
Seven fields. And why logs don't count. What your risk team will ask for, and what most systems can't produce.

What is an agent approval policy?
An approval policy decides which agent actions go ahead, which wait for a person, and which never happen. Most teams have a paragraph. You need a table.
How to monitor AI agents in production
Uptime tells you it's running. Not that it's right. Four signals to watch: quality drift, hold rate, spend per run, latency.