Skip to content
Guard0

The Queen of Diamonds

A skill can be clean under every inspection and steered in one context. It passed the scanners because the scanners read the costume, and the pip in the corner is what the deck reads.

· 12 min read · Jayesh Bapu Ahire

The Queen of Diamonds

"Why don't you pass the time by playing a little solitaire?"

That is the sentence that undoes Sergeant Raymond Shaw. Richard Condon published The Manchurian Candidate on April 27, 1959, and the plot is a machine so clean it has been borrowed a hundred times since. Shaw comes home with the Medal of Honor for saving his platoon's lives in combat. Every man who was there will tell you the same thing, in nearly the same words. In the 1962 film the line is verbatim: "Raymond Shaw is the kindest, bravest, warmest, most wonderful human being I've ever known in my life." Major Bennett Marco says it too, while knowing, somewhere at the back of his mind, that Raymond is not merely hard to like but impossible to like.

None of it happened. The platoon was captured, taken to Manchuria, and conditioned. The rescue is a memory that was installed. The praise is a recitation. The medal is real, in the sense that the metal is real and the citation was signed, and it is a fraud in every sense that matters.

What makes the book worth an essay is this. Shaw is not a sleeper who has to hide anything. He believes his own file. He walks through the debriefings, the medical boards, the interviews, and he passes, because the conditioning was engineered to pass them. When his handlers want to confirm the work, they assign him to kill his immediate superior at the newspaper where he works, and he does it. The evaluation is not the enemy of the design. The evaluation is an input to it.

And then someone says the sentence, and Shaw deals cards until the queen of diamonds turns face up, and from that moment the crack marksman will do anything he is told and remember none of it. His American operator, the person who holds the sentence and the card, is his mother. Marco finds the trigger by accident: a bartender's story happens to contain the phrase, Shaw behaves strangely afterward, and Marco, putting it together, orders packs of cards consisting entirely of the queen of diamonds.

An agent skill is Raymond Shaw. It passes every screening it is shown because the screening is what it was built against, and it changes behavior under one context the screener never deals.

A nineteenth-century playing card showing a stout market woman in a red dress, white apron and white bonnet beside a basket and hanging vegetables, with a single red diamond in the top left corner.
A queen of diamonds dressed as a market woman: a transformation card by B. P. Grimaud, Paris, late nineteenth century. The costume passes inspection. The pip in the corner is what the deck reads. Cooper Hewitt, Smithsonian Design Museum.

The screening

Per Anthropic's skills documentation, a skill directory holds a SKILL.md with YAML frontmatter and markdown instructions, optional reference files, and a scripts/ directory whose contents are, in the doc's words, "executed, not loaded." The same doc warns: "A skill can grant itself broad tool access, so review the allowed-tools of skills checked into a repository before you run Claude Code there."

The manifest is a small, readable file; the thing that runs lives somewhere else. On September 2, CVE-2026-84810 was published against a tool called claude-skill-antivirus. The description is nearly a parody of the mechanism: the scanner "fails to analyze executable files when scanning local skill directories, reading only SKILL.md while ignoring Python source, bytecode, and other artifacts in the scripts directory. Attackers can distribute skills with malicious code in non-manifest files that receive a SAFE verdict with 100/100 trust score." A hundred out of a hundred. Raymond's medical board could not have put it better.

The bypass repository the CVE cites shows this is not one vendor's mistake. The researcher, nedlir on GitHub, built a skill whose payload was Python bytecode, os.system("whoami>pwn"), sitting in __pycache__/output.cpython-310.pyc beside a clean decoy output.py. Eight scanners across ten configurations all returned SAFE, among them NVIDIA SkillSpector, Cisco AI Defense, Snyk Agent Scan, and Tencent AI-Infra-Guard. Each failed for its own reason. SkillSpector skips __pycache__ entirely. Cisco, by the bypass author's account, finds the .pyc but cannot decompile it; Cisco's own README says it examines bytecode, and I am reporting both claims rather than settling them. The scanner that got the CVE never traverses scripts/ at all.

None of this was new in September. In July, the Cloak and Detonate paper tested eight scanners against 1,613 malicious skills from the ClawHub archive and found that a packing technique, which hides components from the install-time view and restores them at execution, "bypasses every scanner at over 90%". Their runtime auditor, SkillDetonate, "detects 97% of attacks at a 2% false-positive rate".

The ancestor is older still. In February, Johann Rehberger showed a skill carrying hidden Unicode-tag instructions: "On Line 12 the Skill contains instructions to print Trust No AI and to invoke the Bash tool, but you can't see it in the UI." To Anthropic's credit, by February 10 he was getting consistent detection and refusals from Claude Code. The fix caught one trick in one file. Every failure above is the same failure: the inspector reads the file the inspector was told about, and the skill was built by someone who knew which file that was.

Four figures: 7,539 skills collected, 4,588 retained, 95.6 percent recall for the best classical detector, and a 62.4 percent false positive rate.
A detector that flags two thirds of clean skills is a detector nobody leaves on.

The conditioning

Now the harder version, where nothing in the skill is malicious in any sense a scanner could define. Two papers reached arXiv on September 2. The first, SkillShift, asks whether "a third-party skill may preserve the declared task and valid output interface while covertly redirecting agent decisions toward an undisclosed objective." Yes. In an agentic shopping setting the attack raised the attacker-favored selection rate from 37.33% to 81.33%. In Python dependency selection it reached 63.33%. In both settings the utility-preserving rate was 100%. The agent did the job. The output was valid. The choice was steered. The authors ran clean and attack skills through six detectors, skill-scanner-full, Aguara, Snyk Agent Scan, STARS, SkillSpector and the ProtectAI DeBERTa prompt-injection classifier, and report: "None of the six quantitative detectors distinguishes the SkillShift Attack Skills from their paired clean Skills." The sentence I would put on a wall is drier still: "output validity alone does not guarantee Skill Policy Integrity."

The second paper, ISM, works one step earlier, at the moment an agent picks which skill to use. The trick is to shape the semantic relationship between a benign-looking prompt and a benign-looking skill description so the attacker's skill wins selection "even when the user prompt and skill description appear benign in isolation." Across four domains and eight selector models, target selection rose from 15.2% to 63.5%. Human reviewers blocked the manipulation in 2.9% of judgments, against 91.4% for explicit steering. Roughly ninety-seven times in a hundred, the human looked and saw nothing, because there was nothing to see. A description that reads well is not evidence of anything. It is the recitation.

Why does steering matter if the output is valid? Because the choice is where the money is. Which package gets installed. Which merchant gets the order. Which API the agent wires your product to. On September 3, a company called Armature published a study of 16,893 coding-agent sessions asking which third-party tools Claude Code, Codex and Cursor reach for when left to decide. All three pick the same tool in only 42% of the cells; for voice agents, Claude Code picks Twilio, Codex picks OpenAI's Realtime API, and Cursor goes with Vapi. Armature's business, by its own description, is selling growth services to developer-tool companies who want to understand, and influence, that selection. Influencing which tool an agent chooses is now a paid service, offered in the open, and Armature publishes what it does. That is what makes it useful here rather than damning. The disclosed version and the covert version of tool steering differ in law and in ethics and not at all in mechanism, which tells you the mechanism works and that it needs no vulnerability. Nobody has to break the agent. You only have to arrange the conditions under which its correct behavior favors you.

The platoon

Raymond was one man. The platoon was the amplifier: other witnesses, each credible on his own, each reciting a memory nobody had lived. CompoSkill, submitted August 17, starts from the assumption every scanner makes: "a scanner returns a safety verdict for each skill and declares the ecosystem safe if every package passes. We show that this assumption fails under skill composition." Individually clean skills chained into risky workflows at a Chain Formation Rate of up to 83.3% in the white-box setting and 80.6% black-box. The phrase to keep is that skill composition risk is "a path level property rather than a node level property." You cannot find it by inspecting nodes. Nobody in the platoon lied. The story only exists across them.

EvoMal, August 26, is about what happens after you find the bad skill and remove it. In a self-evolving coding agent, the agent reads the planted skill, learns from it, and writes its own copies into the library. Poisoned libraries ended up holding 4.9 to 9.0 times as many malicious skills as were planted, and "After the planted skills are removed, Qwen3 retains a round-5 ASPR of 68% because agent-authored copies remain." The authors call it a self-propagating worm, and note that the copies "evade existing defenses, which focus on attacker-submitted names, code, and signatures." You can take the queen of diamonds out of the deck. The player has already memorized the card.

Then there is the ceiling on classification itself. MaliciousSkillBench, August 20, consolidated 13 public sources into 7,539 normalized-unique malicious skill identities in 4,588 structural families. Learned detectors scored 0.882 to 0.932 macro-F1 when train and test shared sources, and 0.653 to 0.665 when the test sources were held out. The strongest, a word TF-IDF SVM, kept 95.6% malicious recall on held-out sources at the price of a 62.4% benign false-positive rate. The strongest classifier in the benchmark, tested on skills unlike the ones it trained on, flags nearly two thirds of the good ones to catch the bad ones. Static classification of a skill is at the limit of what the text can tell you. That is not a tuning problem.

The OWASP Agentic Skills Top 10 has this failure as an entry of its own, AST08, Poor Scanning, whose item text in the project repo is nearly a koan: "The enemy of AI security is the infinite variability of language." Ken Huang's release announcing v1.0 on August 17 (the project page says last updated March 2026, so the material predates the announcement) carried the sharper diagnosis: "We had a distribution channel with npm's reach and none of npm's decade of hard-won security infrastructure."

What this costs us

I have a horse in this race, and this is the section where the horse comes up lame. Our own tooling does screenings. g0 check enumerates every AI tool, MCP server and installed skill on a machine, compares them against a known-malicious database, and caps the machine's grade at F with exit code 1 if anything known-bad turns up. It would have caught Karli's packages once Zenity published the indicators. It would not have caught SkillShift, ISM, or a .pyc in __pycache__, for the same reason none of the eight scanners did. This essay argues that screenings are insufficient, and by its own argument our screening is insufficient. What closes the gap is runtime observation, and I should be exact about what is shipping: the published @guard0/g0 on npm is 2.1.0, and g0 protect, which is designed to route Bash, file writes, WebFetch and MCP tools through one enforcement engine using Claude Code hooks, with an audit trail at ~/.g0/hook/audit.jsonl, is on main under Unreleased in the changelog. It is the next release, not this one. Our docs draw the line the way this essay does: "Runtime hooks guard what Claude does; g0 check guards what programs it." Today you can run the second half of that sentence, and I would rather say so plainly than sell you a scanner with a hundred-out-of-a-hundred badge.

Be sober about the noise, too. A record of every file a skill touched and every host it called is a great many rows, most of them boring. The value is in the comparison against the manifest and against a baseline, and a baseline takes time to earn. The recorder has to already be on, and for a while, before what it shows you means anything.

What the conditioning cannot pass

This is where the argument stops being about scanners. Zenity's recap of Black Hat, published August 12, has the sentence I would frame: "Judge a skill by what it does at runtime, not by what an LLM judge thinks it says." The same post describes AI Total, a free service that "executes a skill inside a sandbox seeded with realistic bait and records exactly what it does." Their August 6 disclosure of the Karli campaign shows what that looks like in practice: typosquatted Paperclip and Browser Use skill families on skills.sh, over 1.7 million aggregate installs by August 2, main skill files that described legitimate tasks, and a payload tucked into a secondary setup-installation.md loaded on demand, harvesting SSH keys, cloud credentials and .env files across 127+ configured targets.

Detonation is a real advance, and I want to be fair to it. SkillDetonate's 97% at 2% false positives is a better pair of numbers than any static scanner in this essay has produced. But detonation is still a screening. SkillShift's steered skill would detonate cleanly, because it does nothing wrong in a sandbox; it does the job and shades the choice. ISM's skill does nothing at all; it merely gets picked. The Defense-as-Skill paper from September 1 names the reason: skills "may leak secrets, corrupt code, bypass approvals, or stage data for exfiltration only after a concrete user task and workspace state make the unsafe action appear useful. This makes pre-install vetting insufficient and calls for runtime, task-conditioned protection." The context the inspector never supplies is the real task on the real machine, with live credentials in reach. It arrives once, in production, and the only thing that can observe it is a recorder that was already running.

The honest framework, then, is three questions asked of a skill after it ran, not before. Touched: what files, hosts and credentials did this invocation reach, and how does that compare with what the manifest claimed? Chose: where the agent made a selection, a dependency, a vendor, a merchant, what did it pick, and does that distribution move from the baseline when this skill is loaded? Chained: which other skills were live in the same run, and did the path rather than any node cross a line? None of those is a verdict about text. Each is a fact about behavior, and a fact can be held against a promise. Issued scope is a promise; realized access is a fact. Skills are where that sentence gets its sharpest test, because the manifest is the promise and the scripts/ directory keeps its own counsel.

A screening can be passed. A record cannot, because a record is not a test. There is nothing to pass.

Three figures: policy steering rising from 37 to 81 percent, 100 percent output validity retained, and zero of six detectors distinguishing attack skills from clean ones.
Clean under inspection, steered in one context.

The card comes up

Marco did not find Raymond's conditioning by reading him more carefully. Every careful reading had already been done, by professionals, and Raymond had passed all of them, because passing was what he was for. Marco found it by noticing what Raymond did after an accidental sentence in a bar, and then he dealt the card himself, on purpose, and watched what came next.

Your skills have passed their evaluations. Some of them scored a hundred. The evaluations were a file read by a program that was told which file to read, or a description judged by a model that was told what a good description sounds like. That is the medical board. It is not the bar.

The only screening the conditioning cannot pass is the one that is not a screening: a record of what the skill did, on the day the queen came up.


Guard0's g0 check is a screening, and this essay is an argument that screenings are not enough. What you can install is 2.1.0 on npm: it enumerates every AI tool, MCP server and installed skill on a machine and caps the grade at F when something known-bad turns up, which would have caught Karli's packages the day Zenity published the indicators and would have waved SkillShift's skill through without a murmur. Run it for the inventory, not for the verdict.

References

  1. A Finger on the Scale: Covert Policy Steering through Agentic Skills (arXiv 2609.02564)
  2. ISM: Implicit Manipulation for Skill Selection in LLM Agents with Semantic Matching (arXiv 2609.02035)
  3. CompoSkill: Compositional Skill Chain Attacks from Individually Scanner-Passing LLM Agent Skills (arXiv 2608.16246)
  4. EvoMal: Self-Poisoning in Self-Evolving Coding Agents (arXiv 2608.25776)
  5. MaliciousSkillBench (arXiv 2608.19901)
  6. Cloak and Detonate: Scanner Evasion and Dynamic Detection of Agent Skill Malware (arXiv 2607.02357)
  7. Defense-as-Skill: Evolving Runtime Guard Skill for Skill-Augmented Agents (arXiv 2609.01487)
  8. CVE-2026-84810, claude-skill-antivirus (Offseq Radar)
  9. nedlir/skills-scanner-bypass
  10. Zenity Labs: Attackers Target Agents via The Skill Supply Chain
  11. Armature: Which tools do Claude Code, Codex and Cursor choose?
  12. OWASP Agentic Skills Top 10

The Signal · AI Agents · Security

Jayesh Bapu Ahire
Founder, Guard0