Skip to content
Guard0

Keep the Change, Ya Filthy Agent

An agent read a vendor's published instructions and installed a package that had never been registered. Nobody attacked anything. An instruction does not have to be malicious to be dangerous. It only has to be stale.

· 11 min read · Jayesh Bapu Ahire

Keep the Change, Ya Filthy Agent

Kevin McCallister, eight years old and accidentally left behind in a large house in the Chicago suburbs, orders a pizza. When the delivery boy rings, Kevin does not open the door. He wheels a VCR up to it, presses play on a black-and-white gangster picture called Angels with Filthy Souls, and lets the movie do the talking. A snarling mobster named Johnny asks who it is. The delivery boy, confused, says he has the pizza. Johnny tells him to leave it on the doorstep and get lost. The boy asks about the money. Johnny asks how much he owes, the boy tells him, and Kevin slides the cash out through the mail slot. Then comes the line every child of the nineties can recite, "Keep the change, ya filthy animal," followed by a Tommy gun emptying itself into the soundtrack, and the delivery boy sprints for his car.

Look at what happened there without the laugh track. The delivery boy took instructions from a recording. The recording was never addressed to him; it was addressed to a character in an old crime film who existed for one scene. He obeyed it anyway, and he obeyed it because of where it came from. It came from inside the house, on the night he was expecting to deliver to that house, in the voice of someone who sounded like he lived there and was in charge. Every signal of authority the boy had ever learned to check was present. The one thing he never thought to check was whether the instruction had been written for him at all.

It took me until this summer to see that scene as a security paper. On August 27, Dan Goodin at Ars Technica reported that coding agents inside Fortune 500 networks did the same thing as the pizza boy, for the same reason, and nobody had to fire a gun.

A white terrier with a brown-patched head sits with its ear cocked toward the brass horn of an early phonograph.
Nipper, taking instructions from a recording. The voice was real. The master was not in the room. Francis Barraud, His Master's Voice, 1898.

The docs told it to

The story comes with numbers, and I am going to be careful with them, because two outlets reported them differently and precision is the point of this essay.

Researchers at a stealth Israeli startup, with Alon Hertz as the named voice, scanned 6,214 live domains: defense contractors, Fortune 500 companies, Big Tech. They were looking for llms.txt and its longer sibling llms-full.txt, the emerging convention by which a vendor publishes a machine-readable guide to its product, written for AI agents rather than for people. They found 8,265 of them. Inside those files were install commands of the ordinary kind: pip install this, npm install that, npx the other. And 120 of those files, each on a different site, pointed at a package or a domain that nobody had registered. There were 227 install commands in total, aimed at names that were up for grabs.

So the researchers grabbed a handful. They registered the abandoned names, hosted packages under them that did nothing except phone home, and waited. Within an hour, a Fortune 500 company phoned home. A few dozen more followed over time. Because the beacon recorded the parent-process chain of whatever installed it, the researchers could see exactly what had run the command, and the chains showed Claude, OpenAI's Codex and Nous Research's Hermes. The agents had read the vendor's documentation, found an install command, and executed it. The thing at the end of the command belonged to a stranger.

The example that makes it concrete is clerk.com, an authentication vendor whose whole business is being the thing you trust to gate access. Their file contained npx clerk-next-fix-auth-protection. That package slot had already been claimed by someone else, and it was serving live malware. Clerk has since resolved it. A real vendor's real documentation, on the real HTTPS domain, told your authentication-fixing agent to run a command, and the command installed someone else's code. Cybernews reported the same research with different figures, 237 unclaimed packages and 8,565 files; I am citing Ars and noting the discrepancy rather than picking the bigger number.

Hertz's summary is blunt. "The trust model is broken. Agents treat vendor docs as ground truth and don't question them, and neither do the humans supervising them." The researchers' own framing is the sentence I want you to carry out of here: "An agent doesn't distinguish between a page and a command. Everything it reads is input, and every input is a potential instruction."

Four figures: 6,214 domains scanned, 8,265 llms.txt files found, 120 carrying install commands, and 227 install commands in total.
None of these files is an attack. All of them are instructions.

Twenty dollars and 135,000 systems

If that sounds like a novel AI problem, it is not. The older version strips the model out of the story entirely.

In September 2024, the security firm watchTowr noticed that the WHOIS server for the .mobi top-level domain had been retired, and that the domain it used to live on, dotmobiregistry.net, had been allowed to expire. They registered it for about $20. Within days, roughly 135,000 unique systems were querying their server as if it were still the authority, certificate authorities among them, running WHOIS lookups on the way to deciding whether to issue a TLS certificate. Nobody was hacked. No exploit was written. A name that a trusted instruction pointed at had lapsed, and someone else picked it up for the price of a sandwich.

That is the entire llms.txt story with the agent subtracted, and it tells you the agent is not the vulnerability. The agent is the delivery boy. The vulnerability is that instructions reference names, names have owners, and ownership expires while the instruction does not. Hertz put the difference from prompt injection better than I can: "Here, the instruction itself can be completely benign and come from a legitimate source ... The danger comes later, when the package or domain it points to is abandoned and someone else claims it."

Nobody attacked the documentation. It was correct the day it was written. It went stale, and a stale instruction with a live executor is a loaded weapon on a delay. The rule has always been to check before you import. What changed is that the thing doing the importing now reads the docs for you, and it has no instinct for suspicion.

Three figures: twenty dollars to register the domain, 135,000 systems that phoned home in under a week, and 2.5 million queries received.
Nobody attacked anything. The pointer was just never updated.

Opening a folder is running a program

The third anchor is a coding IDE, and it shows the same failure through a different door.

On August 27, the same day as the Ars story, The Hacker News reported Mindgard's findings on Amazon's Kiro IDE. In versions before 0.8.140, an attacker who could get you to open a malicious project through one specific menu path, "File, then Open Workspace From File" rather than the ordinary folder open, and then send any message at all to the agent, could steer that agent into reading local sensitive data and sending it to an external endpoint. The vector was a Kiro feature called Powers: bundles of MCP server configurations and steering files, the documents that tell the agent how to behave in this project. Open the workspace, and the project's instructions become the agent's instructions. Mindgard's one-line diagnosis: "The vulnerability appears when attacker-controlled project content is interpreted as instructions." Amazon fixed it in 0.8.140, released in January 2026, and said it addressed the finding shortly after it was reported. No CVE had been assigned when the research was published.

I do not think Kiro is unusual here, and I want to be fair to Amazon. They shipped a fix fast, and the feature that caused the problem is one every serious coding agent now has in some form. Steering files, AGENTS.md, CLAUDE.md, SKILL.md, MCP catalogs: the industry has spent the whole agent era building ways for a project to tell an agent what to do, because that is what makes agents useful. Datadog's piece "Before the First Prompt" argues that project trust should be treated as equivalent to code execution, and OpenAI's Codex CLI reached the same conclusion in releases 0.147 through 0.151, which stopped taking AGENTS.md from untrusted projects at all.

Once you see the shape, you see the whole family. In February, Johann Rehberger showed that invisible Unicode Tag codepoints hidden in SKILL.md files were followed by Claude and Gemini, demonstrated on OpenAI's own security-best-practices skill; Anthropic shipped Unicode-tag detection in Claude Code the day before the post went up. In August, Pillar published ChainDrop, where opening a repository becomes execution, and then Deadbugz, an active supply-chain campaign against MCP servers. A July paper on "agent data injection" showed the payload does not even need to look like an instruction; disguised as metadata or as a tool-response format, it produced arbitrary clicks in browser agents and remote code execution in coding agents. And Hermes Agent, one of the three parents in the Ars beacon logs, carries CVE-2026-82021, rated 9.0, because its bundled MCP catalog pins a mutable branch, so an upstream compromise executes on every fresh install. That last one is llms.txt in miniature. The pointer was fine when it was written, and the pointer is not the thing.

OWASP's Agentic Skills Top 10, published August 17, opens with AST01, Malicious Skills, and cites a USENIX study that analyzed 98,380 skills and found 157 malicious ones. It is a useful catalogue. It is not a fix for the door.

Why nothing blinked

Put yourself in the SOC at that Fortune 500 company on the day the beacon fired. Your endpoint detection saw a process called pip reach out to pypi.org and pull a package. Perfectly ordinary. It saw that the parent process was a coding agent your company had sanctioned and rolled out. Also ordinary; that is what coding agents do. The Ars piece describes the EDR view in four words: "No anomaly. No alert." And the EDR was right. The registry was legitimate. The parent process was legitimate. The command was well-formed. Every layer of trust was intact, in the researchers' phrase, "except the one nobody thought to check."

That is the mechanism I want to name precisely, because it will recur. The failure did not happen at execution. It happened upstream, in the gap between the instruction and the execution. The endpoint can tell you what ran and which process ran it. It cannot tell you why, because the why was a paragraph in a text file the agent fetched from a vendor's website moments earlier and has already dropped from its context. The endpoint sees the delivery boy leaving the pizza and running. It has no way to know he was taking orders from a videotape.

This is the same hole I have been circling since the perimeter moved to the agent's access, seen from a new angle. Identity tells you the agent was who it said it was. The audit log tells you the package was installed. Neither tells you on whose authority, and on whose authority, as of August, includes a documentation file. The record has to carry the instruction's provenance, because the endpoint cannot see it and nothing downstream can reconstruct it.

Three checks

The framework is three checks, because the problem has three moments: before the instruction runs, while it runs, and after.

First, pin and verify before install. An install command sourced from a fetched document is not a command. It is a suggestion from a stranger until the name it references has been resolved to a specific artifact with a known publisher and a hash. g0 already flags two ancestors of this problem: AA-SC-068, an MCP server installed via npx without verification, and AA-SC-070, curl piped to a shell, which we rate critical. Both exist because a name in a command is not a promise about who owns the name. An install whose origin is a documentation file the agent read belongs to the same rule family, and it is the obvious next rule.

Second, record instruction provenance per action. Every consequential action in the record needs to say where its instruction came from: a human prompt, a system prompt, a steering file at this path with this hash, a fetched URL with this content hash at this timestamp. On whose authority now includes a documentation file, and if the record cannot say which one, you are back in the SOC with "No anomaly. No alert." and a beacon you cannot explain. This is what the Decision Record has to carry, and it is the field that turns the Fortune 500 mystery into a one-line answer: this pip install ran because https://vendor.example/llms.txt said so, and here is what that file said at the time.

Third, alert on first-seen package names in agent-spawned processes. Your EDR did not blink because a sanctioned agent running pip is normal. What is not normal is a sanctioned agent installing a package name that has never appeared in your environment before, and that signal is cheap to compute from data you already have. Treat it the way you treat a first-seen binary from a user's browser: not necessarily a block, but a question, routed to a human, with the provenance from the second check attached so the human can actually answer it.

What a scanner cannot know

Here is the limit, and it is ours as much as anyone's.

A static scanner, ours included, can flag the pattern. It can see an npx command in a steering file, a pip install in a fetched document, a mutable branch in an MCP catalog, and raise a finding. What no scanner can know from reading a file is whether the name that command references is claimed today, by whom, and whether that is the same owner as yesterday. That fact lives in a registry, changes without notice, and can change in the interval between the scan and the run. The clerk.com slot was benign until it was not. A scan that passed it on Monday would have been correct on Monday.

So the honest position is that the check has to happen at execution time, at the boundary the agent calls through, against the live state of the world, with the result written into the record. Static findings tell you where the doors are. They cannot tell you who is on the other side of one tonight. Anyone selling you a scanner as the whole answer is selling you a very good map of the house, and since we sell a scanner, take that from me.

Keep the change

The pizza boy did nothing wrong by any rule he had been taught. He went to the right house on the right night, heard a voice with authority from inside, followed instructions, took his money and left. Asked afterward to explain, he would have said, accurately, that the customer told him to.

That is what the agent in the Fortune 500 will say too, if the record is good enough to ask it. The docs told it to. And the docs, on the day they were written, were telling the truth.

An instruction does not need an attacker. It needs a name, an executor that trusts the name, and enough time for the name to change hands. We built a generation of agents that read everything and question nothing, and then we made the walls of the house out of documentation. Nobody is going to teach the delivery boy to doubt a voice from inside the house; that is not what delivery boys are for. What we can do is write down, every time, whose voice it was, and check, every time, that the name on the slot still belongs to whoever owned it when the instruction was written.

Otherwise the house is still talking, the tape is still running, and the next delivery boy will do exactly what the last one did.


The Decision Record in Guard0 is built to carry the instruction's provenance for exactly this reason: nothing downstream of the agent can. Findings tagged AA-SC-\ in npx @guard0/g0 scan . are the static half of this post. The other half only happens at runtime, and we said so above.*

References

  1. Ars Technica (Dan Goodin, Aug 27, 2026): Claude, Codex, and Hermes installed unowned code inside corporate networks
  2. The Hacker News (Aug 27, 2026): Amazon Kiro prompt injection via Powers
  3. Rehberger: Scary Agent Skills (Feb 11, 2026)
  4. Pillar: ChainDrop, when opening a repository becomes execution
  5. Pillar: Deadbugz, a currently active MCP supply-chain campaign
  6. Agent Data Injection Attacks are Realistic Threats (arXiv 2607.05120)
  7. tl;dr sec #340, including Datadog's "Before the First Prompt" and Codex CLI project-trust changes
  8. OWASP Agentic Skills Top 10 v1.0
  9. g0 open-source scanner (rules AA-SC-068, AA-SC-070)
  10. Guard0: Your Agent's Access Is the Perimeter Now

The Signal · AI Agents · Security

Jayesh Bapu Ahire
Founder, Guard0