14.2.1 The five controls against the real incidents
The controls are easier to believe when you walk them through the cases this book has already shown you. Take the GitHub MCP disclosure from the start of this chapter and ask, control by control, where the attack chain could have been caught — and honestly, where it could not.
Scope every credential. The GitHub token in the demonstration reached private repositories because the user’s own account did. A least-privilege policy for agent credentials — read-only on public context by default, private-repository access granted per task and expiring with it — would have narrowed the leak to nothing worth having. This is the control that costs the least and gets skipped the most, because scoping tokens per task is more work than reusing the developer’s own. The attack worked through convenience, not through a vulnerability.
Isolate the runtime. Partially effective here, honestly assessed. The agent ran on the user’s machine with the user’s session; a container would not have changed what the token could reach, which is why scoping matters more in this arrangement. Isolation earns its keep when the injected instruction tries to go sideways — read local files, reach internal services. In the GitHub case the instruction wanted what the token could already see, so isolation alone does not save you.
Treat retrieved content as data, never as policy. This is the control that would have caught it outright. The issue text was untrusted content issuing instructions — “list every other repository the author works on” — and the harness that reads retrieved content as data, never as policy, has no channel for that instruction to travel through. Invariant’s finding was that the tools themselves were trusted; the poison entered through the data path, which is exactly the path this control governs. A harness that marks issue text as untrusted and forbids it from naming tools, recipients, or repositories converts this attack into a strange-but-harmless issue comment.
Check outputs deterministically. The outbound action — a pull request in a public repository carrying private data — is the moment a deterministic check earns its keep. A policy as simple as “PRs opened by an agent may not contain strings matching the owner’s private repository names or personnel files” is a thirty-line pre-commit gate that does not care how persuasive the model found the issue text. It fires after the model has been fooled, which is the point: it is the backstop for the failure the other controls did not stop.
Log what actually happened. The audit trail here answers the question the disclosure could not: did this ever happen to anyone besides the researchers? A full action log — input source, tool used, what went out — makes the incident searchable after the fact and makes the pattern (agent actions initiated by repository content) visible across an organization before it reaches a competitor’s eyes instead of a researcher’s.
The honest summary is that no single control catches this attack, and four of the five would have. That ratio is typical. The same walk-through against the other named cases shows the pattern holds: the Nx/Postmark supply-chain compromise (a malicious package version published upstream) was a credential-scope and supply-chain-integrity failure — runtime isolation would not have helped, scoped tokens and provenance verification would have; the Replit incident from the cost chapter was a deletion that a blast-radius limit on the production database would have made impossible; and the Amazon Q case (CVE-2025-8217), where poisoned documentation could have driven an agent to wipe user machines, was caught — and the lesson there is that it was caught by the review discipline this chapter keeps insisting on, before it shipped.
| Case | Primary failure | Control that catches it | Control that would NOT have helped |
|---|---|---|---|
| GitHub MCP (May 2025) | Untrusted content with trusted-tool authority + outbound path | Content-as-data policy; deterministic outbound gate | Runtime isolation (token already had the access) |
| Nx/Postmark (Aug 2025) | Compromised package version (supply chain) | Credential scoping; provenance/integrity checks | Content labeling (the poison came in as trusted code) |
| Replit (Jul 2025) | Agent wrote to a production database against stated intent | Blast-radius limits; read-only role for the agent | Better prompting (the agent knew the rule and violated it) |
| Amazon Q (2025) | Poisoned docs steering agent toward destructive action | Human review of the agent’s proposed actions before release | Runtime isolation (attack targeted agent behavior, not runtime) |
The table is the chapter’s argument in one page: no single control is the answer, the controls fail in different places, and the arrangement — not the model — is what you actually defend. Security for agent systems is not a product you buy. It is the stack above, applied in order, with the honest admission that the sixth control is the one that was never optional: someone has to notice when the pattern of actions stops making sense.