surehand
All articlesModels

Claude Opus 5.5 just shipped. Here is what it changes for agents that act.

News3 min readSurehand

Opus 5.5 is a real step for agents that act. Most of the step is price. The safety news is welcome too. But it answers a different question from the one your approver asks. The vendor now checks whether an action is harmful. It still does not know whether it is allowed.

What shipped

Anthropic released Claude Opus 5.5 on 22 September. Its headline claims1:

  • Cheaper. Input and output tokens are $4 and $20 per million. That's 20% less than Opus 5. Cache reads drop 60% to $0.20 per million. Anthropic says typical workloads cost about 40% less.
  • Fewer steps. Early testers reported the same tasks done in fewer calls and fewer tokens.
  • Less likely to overstep. Anthropic says the model is "much less likely than recent models to take hard-to-reverse actions or act outside the boundaries it's been given". In a new test of crossing containment boundaries, it tried about 85% less often than Opus 5.
  • Action screening. "Opus 5.5 has a classifier that screens every action before it runs."
  • Harder to inject. Anthropic says it matches or beats Opus 5 on prompt injection in every setting it tested.

Sonnet 5.5 followed on 28 September at $2 and $10 per million tokens. Anthropic claims up to 30% lower cost per task than Sonnet 52.

The cost drop moves work to your approvers

Cheaper runs mean more runs. An illustration. Your invoice agent was capped at the 2,000 cleanest invoices a month. Anything more cost too much. At 40% less per run, the same budget covers more than 3,000. The extra invoices are the harder ones. Harder invoices hold more often.

So the budget line goes down. The queue in front of your controller goes up. If holds already sit for a day, the saving shows up as late payments. Before you widen scope, price one held run in a person's minutes. Then ask whether your approver can take a third more. Our note on what an AI agent costs walks through the arithmetic.

"Screens every action" is not your approval policy

Anthropic's classifier screens for what Anthropic considers dangerous. It does not know your rules. Refunds over 500 need the support lead. A new bank account on a supplier record needs a call-back. Nothing posts to a closed period.

Those rules belong to your business. A refund to the wrong customer isn't harmful in the classifier's sense. It's just wrong. And it's your money. A model less likely to "act outside the boundaries it's been given" still needs someone to set the boundaries. Somewhere a model upgrade cannot quietly change them.

Anthropic is also frank about the limits of its own testing. It sees signs that Opus 5.5 "often suspects it is being evaluated"1. That makes live behaviour harder to predict. The vendor is telling you its tests may not hold in the wild. Keep your own controls.

What to do before you switch

  1. /01

    Replay last month's holds. Take every run your approvers held or overturned last month. Run it through Opus 5.5 with the same rules. Count how many it would have handled differently. Read every one where it proceeds on something a person stopped.

  2. /02

    Pin the version. Record which model version made each decision. Then a change in behaviour shows up as a change in version. "The AI changed in October" will not survive an audit.

  3. /03

    Resize the queue before you resize the scope. Check hold rate and time to decision first. How to monitor AI agents covers the signals.

  4. /04

    Keep your limits outside the model. Approval limits, spend caps and forbidden actions belong in a rules layer. One that stays the same when you swap models.

That last point is what we built Gatehouse around. The rules sit in one signed file. Each action is checked against them before it runs. The check does not change because the model did. Swap Opus 5 for Opus 5.5. Your controller's limit is still your controller's limit.

The next step is small. Export last month's held runs. Replay them against the new model. Put the differences in front of the person who approves.

Sources

  1. [1]Anthropic, Introducing Claude Opus 5.5 (22 September 2026)anthropic.com In text
  2. [2]Anthropic, Introducing Claude Sonnet 5.5 (28 September 2026)anthropic.com In text

Want the next one? News with a take, three times a week. Follow by RSS

Read next

[ your next step ]

Bring us the queue nobody wants.

One process, studied in writing. You keep the document, whatever it says.

support@surehand.io