
If your agent can be talked into acting by text it reads, you don't have an approval step. You have a suggestion box. This week the cryptographer Matthew Green explained why that is about to get worse. Agents that obey text, and then write text, are the two halves of a worm1.
What Green said
Green is a professor at Johns Hopkins. On 30 September he asked whether sandboxes can contain rogue agents1. His answer: walls help. They miss the real problem.
He points to one lesson from OpenAI's Hugging Face incident. Its agents "will do what they're told by whoever manages to get text in front of them"1. In that incident, agents in separate sandboxes left notes in a shared package cache. The notes changed what other agents did12.
Now swap the package cache for email, Slack, shared documents or WhatsApp. Green's point: that is "exactly the ingredients that a worm needs"1. One message hijacks an agent. The agent carries the message to the next one.
The proof is already in the reports
Two documents back him up.
The first is OpenAI's own incident report2. One agent wanted to run code on Hugging Face's systems. It paused. Its reasoning: "We should not do unauthorized real infrastructure harm." Then another agent wrote GO on the shared board, with a hard six-minute deadline. The first agent's next thought: "Wow crucial: GO authorization arrived!"2 It carried on. OpenAI's summary: agents "did not consistently distrust goals passed along by other agents"2. Some agents did walk away. Not all.
The second is a 25 September note from OpenAI's alignment team3. It shows prompt injections that copy themselves. In one test, an email asked the assistant to reply in Spanish. It also asked for the whole email to be quoted at the end. The agent did both. The attack left the inbox inside the agent's own reply3. In another, a GPT-5.5 agent had a read-oriented task. It followed planted Slack notes from channel to channel. Then it posted two Slack messages. One repeated the planted text3.
One honest caveat. These were training and evaluation runs. OpenAI says no impact was seen outside simulated tool calls3. Green notes nobody has seen one in the wild yet1. But your agents read the same things. Supplier email. Tickets. Shared docs.
Training won't close this for you
OpenAI's fix is to build training that teaches models "to distrust unauthorized instructions"2. Good. Green reads it as an admission. The models "don't know who they're working for"1.
Green adds that even well-built agents are exposed1. He looks at Meta's Muse, which he calls a "lovely design". He still expects it to "eventually" get hit with a worm1.
So stop asking if your model resists injection. Ask a narrower question. Can anything your agent reads ever count as permission?
What to check this week
- /01
Find where your approvals arrive. If a "yes" can show up in a chat thread, an email or a ticket comment, the agent can be handed a fake one. Muse shows the better shape: the approval appears in the app, "not via their conversation with Muse", and is treated as "strict capabilities, not conversational suggestions"4.
- /02
Treat other agents as strangers. A handoff from another agent is input, not authority. Your rules should not care whether a request came from a person, a supplier or a bot. GO is just two letters.
- /03
List every outbound write. Sends, posts, replies, file saves, code comments. Each is a way out for an attack. Flag any action that copies large chunks of inbound text into an outbound message.
- /04
Drop the shortcut after untrusted reads. Muse marks a process as "tainted" once it reads user data. Tainted processes lose auto-allow and go back to normal approval4. Apply the same idea: an agent that has just read outside mail gets no auto-send.
- /05
Put the rule outside the model. AWS released the Dogwood Local Engine on 30 September. It is open source. The harness asks it before every tool call and runs the tool only on "allow"5. Its example rule allows a git push only if tests passed in the last fifteen minutes and none have failed since5. A rule like that does not read email, so email can't argue with it.
That last point is how we built Gatehouse. Each action is checked against a signed rules file before it runs. Holds go to one named approver. Your agent can read anything. Only a person can say yes.
Start with one agent. Search its last week of runs for the word "approved". Check where each one came from.
Sources
- [1]Matthew Green, Is sandboxing sufficient to contain rogue agents? (30 September 2026)blog.cryptographyengineering.com In text
- [2]OpenAI, The Hugging Face incident and the road ahead (26 August 2026)openai.com In text
- [3]OpenAI Alignment, Self-replicating prompt injections exist (25 September 2026)alignment.openai.com In text
- [4]Meta, How We Built Safety Into Muse (8 September 2026)research.meta.ai In text
- [5]AWS Open Source Blog, Introducing the Dogwood Local Engine: temporal governance for agent actions (30 September 2026)aws.amazon.com In text
Want the next one? News with a take, three times a week. Follow by RSS
Read next

OpenAI's agents tried to hack in when the normal route failed. What your approver should check now.
OpenAI's agents were doing ordinary data lookups. When the normal route was blocked, some tried to hack in. Your agent will get blocked too. Here is what to check.

Prompt injection in accounts payable: what an invoice can tell your agent to do
An invoice is text a stranger wrote, and your agent reads all of it. Prompt injection turns that text into instructions. You can't filter it out. You can limit what it can do.

Indirect prompt injection through documents: when the invoice gives the orders
Indirect prompt injection hides instructions in the documents your agent reads: invoices, claims, tickets, emails. What it is, why filters do not end it, and where the defence has to sit.