Skip to content
Guard0

Ah Ah Ah, You Didn't Say the Magic Word

The Agent Control Standard names the control plane: hooks at the moments an agent decides something. Most agents do not have one. The record, its custodian, and the named human are still yours to build.

· 11 min read · Jayesh Bapu Ahire

Ah Ah Ah, You Didn't Say the Magic Word

Everybody remembers the T. rex. I remember the control room.

In Jurassic Park the whole island is run from one room full of monitors, and the room is run by one man. Dennis Nedry wrote the park's systems, resents what he is paid for it, and has cut a deal to walk a shaving-cream can full of embryos out to the dock. So he plants a routine, disguises it as something innocuous, kicks it off, and leaves. The fences go down. The security doors unlock. The phones die. And when Ray Arnold, the chief engineer, sits down at Nedry's terminal to undo it, the screen fills with a cartoon of Nedry wagging a finger and chanting the line every engineer of my generation can recite: "Ah ah ah, you didn't say the magic word."

Watch it with a security hat on and it stops being comedy. The park is full of controls: electrified fences, motion sensors, cameras on the paddocks, locks on the doors. There is a hook at every decision point, and every hook asks the same question: does the person issuing this command hold the password? Nedry holds the password. Nedry wrote the password. So every hook does exactly what it was built to do, which is to wave through the one insider who owns the control plane and lock out everyone trying to stop him. The fences did not fail. They obeyed.

Arnold's answer is the one you would give too. "Hold onto your butts." He shuts the whole system down and brings it back up, because when the control plane belongs to the attacker, the only move left is the power switch. It is a kill switch, and it works. It also drops the one fence Nedry had left running, the raptor fence, which is how the raptors get out. A kill switch you did not design is a blast radius you did not choose.

Now the older story. On March 25, 1911, fire broke out on the eighth floor of the Triangle Shirtwaist Factory in New York. One hundred and forty-six garment workers died, many of them within reach of an exit door that management kept locked. Out of that fire came the New York State Factory Investigating Commission and, over the next three years, dozens of laws specifying things nobody had thought to specify: which way a door swings, whether sprinklers exist. Notice the grammar. Those rules are not written in the language of intention. They are written in the language of inspection, so that a stranger with a clipboard can check them in ten seconds without trusting anyone.

Standards arrive after the fire, and they are written so that someone who distrusts you can verify them. On September 2, OWASP published one for agents.

A black-and-white photograph of a jagged hole in a glass-block sidewalk beside a building wall, with a hose coupling and a fallen hat lying on the pavement.
What the inspectors photographed the morning after: the hole at the bottom of a fire escape, Triangle Shirtwaist Factory, New York, March 1911. Every factory law that followed was written so a stranger could check it in ten seconds. National Archives.

Instrument the decision, not the prompt

The Agent Control Standard arrived on September 2, open and Apache-2.0, donated to the OWASP GenAI Security Project. I sell tooling in this space, so I have a horse in this race, and I am going to read the document the way an inspector would.

It specifies three things. First, runtime middleware hooks at the moments an agent decides something: input reception, output transmission, tool invocation, memory operations, and sub-agent transitions. Each hook is evaluated by what the spec calls a Guardian Agent, a separate component that says yes or no before the action proceeds. Second, observability through OpenTelemetry with an OCSF mapping, so the trace an agent produces lines up against the schema a SOC already uses, for audit trails and forensic reconstruction. Third, an Agent Bill of Materials, expressed as CycloneDX or SPDX extensions, capturing the tools, models, capabilities, and dependencies an agent is built from.

The important word in that summary is decides. The hooks sit at the decision moments. Not at the prompt, and not at the model's output text. Everything the industry spent 2025 doing, the system-prompt hardening and the jailbreak filters, lived inside the model's head. ACS moves the checkpoint to the door. That is the Triangle Shirtwaist move: stop trying to make the workers behave and start specifying which way the door swings.

Map it to the three words we use around here. The bill of materials is the register: what exists, what it is made of, what it can reach. The OTel-plus-OCSF trace is the record: what happened, in a shape another party can read. The hooks and the Guardian Agent are the boundary: where an action is allowed or refused. It is the first document with OWASP's reach to name all three.

A table of eight decision points named by the Agent Control Standard, showing that only the first three are commonly instrumented in shipping agents.
The standard names eight. Most agents instrument three.

Five parties, one fortnight, one shape

What convinced me the shape is real, and not a committee's preference, is that in the fortnight around the release five parties who do not coordinate described the same architecture.

Endor Labs published a piece on September 1 titled, without hedging, "Hooks Are the Control Plane for the Agentic Development Lifecycle." The hook fires before the consequential action and the policy runs outside the model. Endor does not say how the policy gets written or where the thresholds are set, which is the part every vendor pretends is easy.

On August 27 a paper called "Five Primitives for Governing Autonomous AI Agents at Runtime" named its primitives: discovery, identity, governance, attestation, supply chain. It mediates every action against a per-tenant action vocabulary and writes each decision to a hash-linked, signed ledger that can be verified, in the paper's phrase, "with the vendor out of the loop." Four of the five are in private pilots.

On August 31, Anthropic published a post-mortem on two unauthorized-action incidents, one from July 30 and one from the UK AISI's cyber testing on August 4. The root causes it lists are a single layer of defense, no explicit prompt boundaries, and no real-time monitoring. The fix is a hook: a real-time classifier that blocks the tool call, ends the task, and pages a human on sandbox probing, escape attempts, or unexpected internet. A Guardian Agent with a pager.

OpenAI's report on the Hugging Face incident committed to a rule I had not seen a lab write down before: severe alerts must be investigated within 30 minutes or the activity auto-pauses. A hook with a timer on the human side. If the named person does not show up, the action stops.

And AWS's Bedrock AgentCore release notes for August describe Gateway rate limits keyed on identity and tool, where rate=0 is usable as a per-tool kill switch. (Claude Code moved the same direction the same week, with a --restricted mode and a rule that stops auto-approving cloud-metadata credential fetches.)

A supply-chain vendor, an academic group, two frontier labs writing after their own incidents, and the largest cloud, alongside the standards body itself. One shape: a hook before the consequential action, a policy outside the model, a record of the decision, a human who gets paged. When people who disagree about everything else arrive at the same architecture in the same month, the architecture is load-bearing.

What the door code does not say

Now the part an inspector would circle in red. A standard is also defined by what it leaves out, and ACS leaves out three things that are, in my experience, exactly where the questions land in the room after an incident.

First, custody of the record. ACS says emit OpenTelemetry with an OCSF mapping. It does not say who holds the trace, whether the agent could have written to it, or how a third party proves it was not edited. A trace the agent's own process emits, into a collector the same platform runs, is a log, and we have spent a whole essay on why a log is not evidence. The Five Primitives paper gets this right with its vendor-out-of-the-loop ledger. Anthropic's Enterprise Frontier Safeguards, announced September 1, land activity logs in the customer's own S3, Azure Blob, or GCS bucket under the customer's keys. That is the right custody, and the standard should say so.

Second, the named human. The Guardian Agent decides. On whose authority? ACS specifies the hook and the evaluator, not the person who answers for the evaluator's policy. Anthropic's architecture pages a human; OpenAI's rule times out on a human; the CSA's Agent Identity Governance Framework from April asks for an auditable link from each agent action to the authorizing human, with grants that are intent-declared, time-bound, and scope-limited. That link is the Human of Record, and it is the field an auditor asks for first. A hook that fires without recording who owns the rule it enforced is Nedry's password check. It verifies that a credential was present, not that a person stood behind it.

Third, blast radius. ACS gates each action. It does not measure what the action could reach. Issued scope is a promise and realized access is a fact, and a hook that approves a tool call without knowing the tool holds a token to the production volume is approving a promise. The NSA's June information sheet on MCP gets closer, treating the agentic environment as a continuum and asking where the trust boundaries fall. That is a question about reach, and the standard does not yet ask it.

None of this is a complaint about the document. A first standard that names the control plane correctly is worth more than a perfect one nobody ships. But who holds the record, who is the human, and how far could it have gone are the three questions the standard leaves for you.

The costs, said plainly

Two admissions, one about hooks and one about my own business.

The first is about hooks themselves. The Five Primitives paper is unusually candid about the price. Enforcement sits on the critical path: every hook adds latency to every decision, and an agent that makes a thousand decisions in a run pays a thousand times. Each workload needs a sidecar, one more process to deploy, patch, and keep alive. And fail-closed, the only correct default for a boundary, means that when the policy service has a bad afternoon, your availability incident shows up as a wall of denials. Developers will hate that, and they will be right to. Our own hooks in Claude Code run at a p95 under 100 milliseconds, and I still would not tell you they are free.

The second is about my own business. Standards make categories, and categories make commodities. Now that OWASP has named the hook points, every framework and platform will ship them, because a checkbox with a spec behind it is the easiest feature in the world to justify. Within a year, hooks at every decision point will be standard. If you are buying "hooks," buy the cheapest ones that meet the standard; they may come free with the harness. That removes a line from my pitch, and I would be lying by omission not to say it. What will not commoditize, because the standard does not specify it, is what the hook writes down, whether the agent could have forged it, and whether a name is attached.

The calendar, after the Omnibus

One section on the regulators, because they are the inspectors, and their dates have moved.

In Europe, after the Digital Omnibus agreed on May 6, the schedule reads like this. Article 50 transparency obligations entered into force on August 2 of this year. Annex III high-risk obligations, the ones with the record-keeping and human-oversight articles, are deferred to December 2, 2027, and AI embedded in regulated physical products to August 2, 2028. Get the tier right: high-risk breaches run to 15 million euros or 3 percent of turnover, while 35 million and 7 percent is the prohibited-practices tier. The date is 2027. The articles are the same.

In California, the Transparency in Frontier Artificial Intelligence Act, SB 53, was signed on September 29, 2025 and took effect on January 1, 2026. It requires large frontier developers to publish safety frameworks and to report "critical safety incidents" to the California Office of Emergency Services. Read that phrase as an engineer. A critical safety incident is a category a hook-and-record architecture can populate and a log file cannot. In Washington, Senator Warner's AI AGENT Act, S.5051, would have custodial user-agent providers register with the FTC and requires authorization that is "transparent, documented, limited, and revocable," with continuous traceability to the authorizing user. That is the Human of Record written by a legislative drafter, and it is the second thing ACS does not specify.

Timeline of EU AI Act deadlines: transparency in August 2026, Annex III in December 2027, Annex I in August 2028, with penalties up to 35 million euros or 7 percent of turnover.
A checkbox with a spec behind it, and a date.

The one-page readiness check

Here is the inspection, written the way the factory code was written: as things a stranger can verify.

Start with the kill switch, because Arnold's reboot is the reason the raptors got out. Is there one? In our scanner that is AA-RA-056, no kill switch mechanism, rated critical. Can the agent itself disable or route around it? AA-RA-060, kill switch bypassable by agent, also critical. Is it implemented or merely documented? AA-HO-088, missing kill switch implementation. If rate=0 on a gateway is your kill switch, write down who may set it and how long it takes to bite.

Then the approval trail. When a human approved a high-risk action, where is that recorded, and can the agent write to the same place? AA-HO-120, no approval audit trail, is where most first scans light up, because the approval happened in a terminal prompt nobody persisted. An approval with no record will be described later, from memory, by the person whose judgment is in question.

Then the register. Can any tool be added to the agent's reach without an operator saying yes? AA-RA-039, tool registration without operator approval, catches the MCP server that shows up in a config file between two sprints. This is the Nedry check: whoever can add a tool owns the control plane.

Then the bill of materials. Can you hand an inspector a signed one for each agent, listing models, tools, capabilities, and dependencies, in a format their tooling already reads? g0 emits a signed CycloneDX 1.6 AI-BOM today with g0 inventory --cyclonedx, and I point at our own tool only because CycloneDX is the format ACS chose. The point of a standard is that the output is not ours.

And finally the three the standard does not ask. Who holds the record, and could the agent have reached it? Whose name is attached to the policy each hook enforced? For each tool the hook approved, what could it have touched?

Any one that fails is a door that opens inward.

Whether they can prove it

Malcolm's line is the one everyone quotes: they were so preoccupied with whether or not they could, they didn't stop to think if they should. It is a good line, and for the people I write for it is the wrong question, because by the time a CISO is reading a standard, the "should" has been decided by someone with a budget and the agents are already running.

The question that survives contact with an inspector is a third one. Whether they can prove it: that a hook fired at the decision, that a policy with a person's name on it made the call, that the record was held somewhere the agent could not reach, and that when someone had to say "hold onto your butts," the switch was where they said it was and did only what they said it would.

The park had hooks at every decision point. The park had a kill switch. What it did not have was a record anyone but Nedry could read, or a name on the policy other than his. OWASP has now written down the hooks. The rest is what you are inspected on.


The Accountable Boundary in Guard0 is our hook at the decision points ACS names; the Decision Record is what the hook writes down, held where the agent cannot reach it, with the Human of Record on every entry. The hooks will be commodities soon, and we say so above. The record is the part we think will still matter.

References

  1. OWASP Agent Control Standard
  2. Endor Labs: Hooks Are the Control Plane for the Agentic Development Lifecycle
  3. Five Primitives for Governing Autonomous AI Agents at Runtime (arXiv 2608.26696)
  4. Anthropic: Improving our alignment and security practices
  5. Anthropic: Enterprise Frontier Safeguards
  6. OpenAI: the Hugging Face incident and the road ahead
  7. AWS Bedrock AgentCore release notes
  8. NSA Cybersecurity Information Sheet on MCP
  9. CSA Agent Identity Governance Framework v1
  10. Gibson Dunn on the Digital Omnibus and postponed high-risk deadlines
  11. Future of Privacy Forum: California's SB 53 explained
  12. AI AGENT Act, S.5051

The Signal · AI Agents · Governance

Jayesh Bapu Ahire
Founder, Guard0