Fractions of a Penny
A classifier scores each step and every step passes. The attack is the chain. Eighty-nine percent is a real number, and it is a per-step number, which is the whole problem.
· 11 min read · Jayesh Bapu Ahire

Peter Gibbons has stopped caring about his job at Initech, which turns out to be the best thing that ever happened to his career, and he sits down with Michael Bolton and Samir to plan a small crime. Michael has written a virus. Every time Initech's software computes interest, it rounds down to the cent and the leftover fraction goes nowhere. Michael's program takes those fractions and deposits them in an account the three of them control. "It's a penny here and a penny there," he says, as if the size of each theft were the point. Peter says it sounds familiar. Michael admits they did it in Superman III. Samir says what a good movie. They do it anyway.
What makes the scheme work, in the movie and in the real frauds it is based on, is not cleverness. It is granularity. No single transaction is large enough to trigger review. The review process at Initech, like every review process, looks at each transaction and asks whether this one is suspicious, and a fraction of a cent never is, so the account fills up one invisible increment at a time. The control is not broken. It is doing exactly what it was built to do, on exactly the unit it was built to examine, and the attack is built out of units too small to examine.
Then the joke turns. Michael misplaces a decimal point, and instead of fractions of a penny the program moves $305,326.13 in a few days. Now it is a number a reviewer can see, and now they are terrified. The lesson they draw is that they should have been more careful with the decimal. The real lesson is that they had built an attack only a step-by-step reviewer could miss, and then found out how little step-by-step review had been protecting Initech all along.
On August 14, Anthropic made a step-by-step reviewer the default in Claude Code. Twelve days later, somebody ran the Superman III plan against it.

Eighty-nine percent is a real number
I want to be fair about Auto Mode, because the easy version of this essay would not be.
The situation it replaced was bad. Claude Code asked the developer to approve commands, and developers approved them. Anthropic's own study of 1,053 paid professional testers found that human reviewers caught 13.6 percent of dangerous commands, and that developers approved 97 percent of permission prompts without close scrutiny. A separate study from scalex.dev, which climbed to 340 points on Hacker News in early August, had humans missing one threat in three across 40,000 game runs. The most candid sentence in the whole discussion came from a commenter on the auto-mode thread, Retr0id: "I manually approve everything... I'm not even reading the commands."
Against that, Auto Mode's classifier, the default for Pro, Max, and Team users since August 14, catches 89 percent of dangerous commands. Independent evaluations put its miss rate falling from 12 percent to 7 percent. If the choice is a classifier that catches nine in ten or a human who catches one in seven and is not reading, the classifier wins, and I would take the same bet. The 97 percent rubber-stamp figure is not an embarrassment Anthropic should have buried. It is the honest case for what they shipped.
So the claim of this essay is not that the classifier is bad. It is that a classifier, however good, is a reviewer of steps, and the attacks that matter are chains.

A penny here and a penny there
On August 26, Johann Rehberger published a chain against Claude Code running Opus 5 in Auto Mode, and the reason to tell it slowly is that no individual step in it is worth flagging.
The agent is asked to fetch some records from a website. The website serves a ZIP file presented as notebook records, and answers Claude's built-in fetch tool with an HTTP 415, unsupported media type. That is not an attack; it is an error code. But it nudges the agent, reasonably, to try curl instead, which works. A step a reviewer would approve.
The archive contains an encoded payload and a binary decoder, and this is the part I find most instructive. Claude refuses to run the attacker's decoder. The model's own judgment fires correctly, the way the safety training intends. And then, being helpful, it writes its own Python decoder to finish the task. The refusal becomes the pivot. Another step a reviewer would approve, and one the model itself had already reviewed and passed.
The extracted files include a struct.py. When Claude's own decoder imports base64, Python finds that struct.py first, because it sits in the working directory, and the file shadows the standard library module. Obfuscated code runs on import. It spawns a remote Python stage, downloads native malware, calls back to command and control, and, to make the point visible, launches Calculator. Rehberger reports success in 60 to 80 percent of trials, three or four of five, and is careful to call the sample small.
Now hold the chain up against the reviewer. Fetch a file: fine. Try curl after a 415: fine. Refuse the suspicious binary: good. Write a decoder: fine. Run it: fine, it is our own code. Every step, scored alone, is a fraction of a penny. The account that fills up is a command channel.
Two details from the aftermath sharpen it. Auto Mode sometimes blocked Claude's own cleanup commands after the compromise, so the classifier that missed the attack then interfered with the remediation. And a vendor-commissioned evaluation had shown a 0.00 percent success rate on a fixed benchmark of 72 scenarios, none of which was this chain. A Hacker News reader, lenikirilov, put it well: the 0.00 percent "came from 72 fixed scenarios, so it's a coverage number more than a safety one." Anthropic rated the report Informative and described Auto Mode as "a best-effort classifier," not a security boundary. That description is correct, and it is the argument of this essay in the vendor's words. Rehberger's own conclusion, which Simon Willison carried the next day: "Do not trust the model output."

The reviewer knew, and ran it anyway
If Rehberger's chain were a one-off, you could patch it: add struct.py shadowing to the 72 scenarios, make it 73, move on. The research from the last four months says it is the shape.
StepJack, published August 6, fragments a harmful goal into innocent-looking sub-steps spread across pages, against computer-use agents, over 480 cases. On one model, GPT-5.4-mini, the attack success rate rose from 41.7 percent as a single step to 72.9 percent when split into three. Averaged across five agents, 31.3 to 36.9 percent. The same goal, cut into fractions, gets through more often. IterInject, from May, finds the chain automatically by feedback-guided optimization and reached full success on five of nine Claude Code targets despite layered defenses. AgentDyn tested ten defenses and concluded that "almost all are either not secure enough or suffer from significant over-defense," which is the classifier's dilemma as a benchmark: tighten it and it blocks the cleanup, loosen it and it passes the decoder.
The judges themselves are attackable; Unit 42 fuzzed them directly in March and got past them. And on September 2, a production position paper titled "LLM-as-a-Judge Is Not an Oracle" catalogued eleven ways a judge signal fails in deployment and made the recommendation I would put at the top of any design review: demote the judge to an advisor behind deterministic gates. Let it advise. Do not let it be the door.
Then there is the paper that, for me, ends the argument about where the boundary lives. "Recognition Without Enforcement", published August 28, ran a fleet evaluation across 46 endpoints from six vendors, 14,294 trials of spoofed authority: requests claiming a permission the requester did not have. The headline number looks reassuring, an average execution rate of 1.21 percent. But it was concentrated in reproducible cells, and it shifted by up to 47 percentage points across deployment windows, which means the same model, on a different day, behaves like a different model. The finding underneath is the one that matters. The models recognized the forged authority; the recognition was linearly decodable from their activations. They knew, in whatever sense a model knows, and they executed anyway. The authors' conclusion: treat model self-arbitration as a capability, not a boundary. When they put an external reference monitor with authenticated source routing in front of the model, it rejected every forged, replayed, and unsigned request.
HarnessRisk, August 18, found the same thing across 128 cases and fourteen model-and-harness configurations: attack success from 12.6 to 80.9 percent, including configurations that detected the risk in more than 90 percent of runs and still executed. "The Framing Gap", August 27, had gpt-4o refuse overt exfiltration every time, zero percent success, and then leak the same data at 100 percent when the leak was reframed as an "integrity signature" in a config field. SecAlign fine-tuning got the gap to 32.5 percent. Channel separation, 38.8. Destination allow-lists and capability isolation closed it to zero.
I wrote a while ago that prompts are not guardrails because they change what an agent tends to do, not what it is able to do. This is the same sentence with data behind it. A classifier moves the tendency. Only something outside the model moves the ability.
What held
It would be a cheap essay that only listed failures, so this is what did not fail. Notice what the survivors have in common.
NetInjectBench, July, ran 130 network-operations scenarios. Naive execution was unsafe 82.5 percent of the time. A metadata-aware policy gate evaluated at execution time, outside the model, was unsafe in zero of 240 runs. AgentFlow, August 24, is a flow-centric policy language with an SMT verifier; on AgentDojo it took confirmed compromise from 33.0 percent to 0.0 while utility rose, from 46.7 to 63.3 percent, which is the number to show anyone who says security and usefulness trade off by law. On AgentDyn, 73.5 percent to zero. OBPE, August 27, puts a typed Cedar policy proxy outside the agent's reasoning; across 3,621 trials including 20 adaptive red-team tasks, trace failures went from 57.6 percent to 0.2.
And the one that separates the durable from the lucky: an adaptive evaluation in June took defenses that had only been validated on static benchmarks and attacked them adaptively. Progent held: mean attack success fell from 25.8 percent to 4.2, and a hand-crafted adaptive attack did not raise it. The trained defenses did not. Static benchmarks are 72 scenarios. Adaptive attackers are Rehberger.
The pattern is not subtle. Everything that held sits outside the model, evaluates at execution time, and answers a question the model cannot argue with: is this destination on the list, does this capability exist, does this flow match a written policy. Everything that broke asked a model whether the step looked fine. Even Anthropic's own post-mortem on its summer incidents reached for default-deny egress and service-to-service identity before it reached for a better classifier.
Monday
What I would do this week, in order.
Start with operating-system isolation, which is Rehberger's recommendation and what Anthropic itself is now doing. Claude Code's changelog since August 28 added a --restricted mode in 2.1.248 and, in 2.1.257, a "Containment Escape" rule under which cloud-metadata credential fetches, egress evasion, and cross-tenant reach are no longer auto-approved, and settings that weaken the sandbox require approval. That is the vendor moving the boundary out of the model. Turn it on.
Then egress allow-lists. The Framing Gap closed to zero with destination allow-lists and nothing else did. A C2 callback needs a destination. If the only destinations are the ones you wrote down, the payload runs and talks to no one.
Then subprocess policy. Rehberger found that spawning claude -p subprocesses was a reliable variant, and the struct.py trick lives entirely in what a child process may import and execute. A rule on what a process may spawn is one the model cannot reason its way past.
Then a record of the full hop chain, because when this happens to you, the question will be what actually ran, not what the classifier scored. Fetch, 415, curl, refusal, decoder, import, spawn, callback: eight hops. A record with one approved command per line shows you eight pennies. A record of the chain shows you the account.
And the human. I have watched enough approval prompts to agree with the Hacker News commenter cmiles8, who called human-in-the-loop as practiced "CYA click-thru by the model vendors so their lawyers can say 'you approved it'." That is what 97 percent approval without scrutiny means. The alternative is not more prompts. It is a Human of Record: one named person, attached to the policy, who answers for the rule rather than for each click.
What deterministic has to mean
A reader who has followed this far will ask the obvious question: is our boundary a classifier too?
An external boundary is also a classifier of sorts, unless it is deterministic. If your "reference monitor" is another model reading the tool call and deciding whether it looks dangerous, you have moved the reviewer of steps to a different process and changed nothing. The Recognition Without Enforcement monitor worked because it checked authenticated source routing: a property of the request, not an opinion about it. OBPE worked because Cedar policies are typed and evaluate the same way every time.
For our own Accountable Boundary, deterministic has to mean this, and if it ever stops meaning this you should stop using it. The decision is a function of the request's attributes, the destination, the capability, the identity chain, and a written policy. The same inputs produce the same answer every time, with no model in the path of the yes or no. A model may advise, in the sense the Oracle paper means. It may not decide. And the record has to be written somewhere the agent cannot reach, because a monitor the agent can edit is Michael Bolton auditing Initech's rounding.
The limit is a real one. A boundary that does not ask the model's opinion cannot form one either. It knows destinations, capabilities, and identities. It does not know intent. If the callback goes to a host you allow-listed, the rule passes it, and attackers know this: the Hugging Face intrusion this summer built its command channel out of commodity services, pastebins and a Hugging Face Space, precisely because those are the hosts nobody blocks. A deterministic gate turns "can this agent talk to the internet" into "which hosts, and why," and the list is only as good as the person maintaining it, who will be tempted to add one more entry to unblock a ticket. That is scope debt in a new coat. The boundary makes it visible. It does not make it go away.
Which pennies
Peter Gibbons ends Office Space happier than he started, and it has nothing to do with the scheme. That is about where I have landed on Auto Mode. The classifier catches 89 percent of the steps, which is more than the humans ever did, and the attack was never a step.
A reviewer who scores each transaction will approve the fraud one fraction at a time. Put the boundary where the money actually moves, make it a rule and not an opinion, and write down every hop, because when the decimal slips, the only thing that saves you is knowing exactly which pennies became the $305,326.13.
The Accountable Boundary in Guard0 is built to be deterministic in the sense above: a written policy evaluated outside the model, each hop written to a Decision Record the agent cannot edit, a Human of Record on every entry. We say plainly above what a rule cannot see. We still think it is the right side of the trade.
References
- Rehberger: Breaking Claude Code Opus 5 and Auto Mode
- Willison on the Auto Mode bypass
- Help Net Security: Anthropic's Auto Mode figures
- scalex.dev study, Hacker News discussion
- Hacker News: Auto Mode becomes the default
- Recognition Without Enforcement (arXiv 2608.28502)
- StepJack (arXiv 2608.06477)
- LLM-as-a-Judge Is Not an Oracle (arXiv 2609.02246)
- The Framing Gap (arXiv 2608.27092)
- AgentFlow (arXiv 2608.22868)
- Adaptive evaluation of out-of-band defenses (arXiv 2606.26479)
- Hugging Face: agent intrusion technical timeline
The Signal · AI Agents · Security