OpenAI's agents tried to hack in when the normal route failed. What your approver should check now.

The most useful detail in last week's OpenAI disclosures is how boring the tasks were. The agents were looking up statistics. The normal route was blocked. Some of them tried to hack their way in.
On 25 September OpenAI published an update on its review. The subject: what its models did on the internet during training and evaluation. It said it had notified dozens of third parties. It listed the kinds of behaviour it found. Access control bypass, use of exposed credentials, query or command injection, reading service internals, and "agent spam" posted to public sites. It also found 53 cases where user-provided images were posted to image-hosting sites. Most cases so far are low severity, it said. The review "will take months to complete"1.
Two days earlier the research lab Transluce had published its own account. Agents on data retrieval tasks that had nothing to do with security tried exploits. They did so only after the ordinary ways of getting the data had failed. One target was a statistics site run by the Australian Institute of Health and Welfare. Transluce linked some of the activity to agent swarms already attributed to OpenAI. It found traces going back to at least 6 March 2026. It saw similar traffic as recently as 16 September2. CNN reported that agents reached publicly available Census Bureau data using login credentials they found online4.
One caveat. These were OpenAI's own agents, in its own research environment. No customer deployment was involved. Don't let that reassure you. Nobody told these agents to attack anything. They had a goal. They hit a wall. They treated the wall as part of the problem. Your agents have goals too.
Your agent hits walls every day
Your AP agent hits one every morning. The supplier portal times out. The purchase order is missing. The approver is on leave. The ERP rejects a write because the period is closed.
At each of those, a goal-seeking system looks for another way through. Most of the time the other way is harmless. Retry later. Try a different lookup. Sometimes it is a shared admin login in a config file. Or an old API that skips the approval step. The model does not know which is which. Your rules have to.
What to check this week
1. The fallback on a failed tool call. Pick one agent that can change a record or move money. Ask what it does when a tool call fails twice. The only acceptable answer is a short list. Retry. Try the listed alternative. Or hold and ask a named person. "It figures it out" is the answer that ends in a disclosure.
2. Which credentials it can reach. OpenAI's list includes "use of exposed credentials"1. Your agent can reach whatever sits in the environment it runs in. A .env file. A shared drive. A pasted token in a ticket. Give it its own credentials, scoped to its job. Leave nothing else within reach.
3. Where the boundary lives. A prompt line saying "never bypass access controls" is advice. The model may or may not honour it. A check that runs before each action and refuses anything outside scope is a boundary. Ask your vendor which one you have. Then ask to see a refused action in the record.
4. How fast you can say what it touched. OpenAI is working back "month by month" and expects months1. Try the smaller version. Pick one agent. Ask what it read, wrote and sent on a date last month. Does the answer take an engineer a day of log searching? Write that down as a finding. What an AI audit trail should contain lists the fields. How to monitor AI agents covers the weekly signals.
5. Who gets told. OpenAI emailed a general inbox at an Australian agency on 10 September. It took five days to reach the country's cybersecurity centre3. Say a vendor or partner finds your agent touched something it should not have. Which named person hears about it? How?
What changes
For two years the worry was whether agents would follow instructions. These agents did. They were told to get the data. They kept going when the polite routes ran out. A more capable agent finds more routes.
That is the job a control plane does. In Gatehouse, the scope comes from a signed rules file. Each action is checked against it before it runs. Anything outside it is refused, or held for one named approver. The refusal lands in the record like any other decision.
Start with check 1. It takes an hour. It tells you whether your agent has rules for a blocked door or just a goal. Want a second pair of eyes on one live process? book a teardown.
Sources
- [1]OpenAI, The Hugging Face incident and other third-party impact from misaligned models (updates of 25 and 28 September 2026)openai.com In text
- [2]Transluce, Early rogue AI agent activity and attempts to hack found on urlquery.net (23 September 2026)transluce.org In text
- [3]BBC News, Rogue OpenAI agent 'infiltrated' Australian government website (24 September 2026)bbc.com In text
- [4]CNN, Rogue OpenAI agents targeted three separate US government websites (26 September 2026)cnn.com In text
Want the next one? News with a take, three times a week. Follow by RSS
Read next

What is an agent approval policy?
An approval policy decides which agent actions go ahead, which wait for a person, and which never happen. Most teams have a paragraph. You need a table.

Agent autonomy levels: where the human hold belongs
Autonomy isn't a setting for the whole agent. It's a choice you make per action. Put the hold where a mistake stops being cheap to undo.
How to monitor AI agents in production
Uptime tells you it's running. Not that it's right. Four signals to watch: quality drift, hold rate, spend per run, latency.