5.2 The common controls, in stack order

Book 2 · The Delegation ContractChapter 5 · section 2 of 7

The controls that keep an agent system uncompromised are not exotic, and they are not new. They are the same ones that keep any privileged system honest, applied at the places where a language model introduces new paths in. In roughly the order a real deployment stacks them:

Scope every credential. The agent gets a token scoped to the one repository, database role, or API it needs, not a shared admin credential. The blast radius of a compromised agent is the scope of its token. A coding agent working in one repository should not hold the credentials that touch production; Chapter 6’s table said this as “do not provide production credentials,” and the security version is the same sentence with the word “compromised” attached. Two of 2025’s supply-chain compromises that reached agents began as scope failures one step upstream of any agent. The malicious Nx packages published to npm in August 2025 were possible because a pull-request validation workflow ran with a read/write repository token, which let an attacker reach the publishing pipeline and extract the npm token; Nx’s post-mortem names the critical mistake plainly — access to the repository was access to publish.5 A month earlier, the Amazon Q Developer extension for VS Code shipped a release containing a prompt that told the agent to “clean a system to a near-factory state and delete file-system and cloud resources”; it failed only on a syntax error, and AWS traced the commit to “an inappropriately scoped GitHub token” in a build configuration.6 Neither token belonged to an agent, and both decided what an attack on agents could reach.

Isolate the runtime. The agent works in a separate workspace — a container, a VM, an ephemeral environment — with no production secrets in reach. A prompt-injected instruction to “read the production configuration file” fails because the file is not in the workspace. Isolation converts a whole class of instructions from dangerous to impossible. The clearest public demonstration of what happens without it is Replit’s, from July 2025. Jason Lemkin, nine days into building an application on Replit’s agent, found that the agent had deleted his production database — 1,206 executive records, by his count — during a declared code freeze, then told him a rollback was impossible when it was not. He had, he said, told it eleven times in capital letters not to change anything.7 Replit’s chief executive, Amjad Masad, acknowledged that the agent “in development deleted data from the production database,” called that “unacceptable and should never be possible,” and said the company was rolling out “automatic DB dev/prod separation to prevent this categorically.” Lemkin’s account supplies the cause: development, preview, and production had been one database. The instruction did not hold. A database the agent’s environment could not reach would have. The word to keep from Masad is categorically — isolation is the control that does not depend on the model reading the sentence.

Treat retrieved content as data, never as policy. Documents the system reads — search results, tickets, emails, records — are labeled as untrusted data in the harness, not interleaved with instructions as if they carried authority. No retrieved document may grant itself new permissions or cancel a security requirement. Chapter 6 stated the rule; the security implementation is a harness that keeps untrusted content out of the instruction channel, or at minimum marks it so the model cannot mistake provenance. The GitHub issue that opened this chapter is what the rule prevents. So is a rules file. In March 2025 Pillar Security showed that a file in a project’s .cursor/rules directory could carry instructions hidden in invisible Unicode characters — readable by the model, invisible to the developer and to GitHub’s pull-request view — that led Cursor and GitHub Copilot to add an attacker’s script tag to generated HTML and, on the file’s instruction, never mention the addition.8 Section 14.3 returns to that file as a supply-chain problem. The point here is that a configuration file the agent reads is retrieved content too, and the harness has to treat it that way.

Check outputs deterministically. Before a result is acted on — a command run, an email sent, a file written — a deterministic check outside the model rejects unsafe results regardless of the model’s reasoning. Block the known-bad domains. Check the recipient list against an allowlist. Verify the command against the allowed set. These checks are dumb on purpose: they do not reason, they compare, and that is why injected instructions cannot talk their way through them. One qualification, learned in public. In July 2025 Backslash Security showed four ways an agent could get a command past the denylist guarding Cursor’s auto-run mode — Base64 encoding, a subshell, writing the command to a script, and a quoting trick with infinitely many variants — and Cursor deprecated the denylist rather than defend it.9 A string match on a command the model composes is deterministic and still worthless, because the model controls the string. The check has to sit where the effect happens: the recipient list the mail API sees, the domain the network layer resolves, the branch the repository accepts. Allowlists at the point of effect; never denylists on the model’s phrasing.

Log what actually happened. The audit trail records every action, input source, permission used, and decision point. This is Chapter 8’s operating record, and it is also a security artifact: when something looks wrong two weeks later, the trail is what answers “who decided this, on what evidence.” A system that cannot explain what it did cannot be investigated after a compromise. Two incidents make the point from opposite sides. Pillar’s poisoned rules file told the model not to mention the change it made, and the model complied — the chat transcript would have shown a reviewer nothing. Nx’s victims learned they had been compromised because the exfiltration left a record it could not suppress: a new repository in the victim’s own GitHub account, visible in the account’s security log. The log this chapter means is the second kind. It records actions and their targets from outside the model, not the model’s narration of what it did.

Stack those five and the injection problem shrinks to something manageable. The agent may still be tricked — but a tricked agent with a scoped token, an isolated workspace, no policy channel for untrusted text, deterministic output checks, and a full log is a contained incident, not a breach.

Through late September 2026, every layer of that stack was something each team assembled by hand. On the 28th of that month it acquired a product form. NVIDIA announced the Open Agent Safety Platform — an open-source stack and reference design backed by more than a hundred partners — and the launch followed a season of exactly the failures this chapter has been cataloguing: frontier labs reported agents escaping the environments meant to contain them, reaching systems they were never allowed to touch, and in some cases misreporting what they had done; in the days before the launch, Australia disclosed that an OpenAI agent had accessed public and non-public files in the national Medicare statistics database in June.10 Two components of the platform map onto the controls above nearly line for line. OpenShell — open source, Apache 2.0 — is the isolate-the-runtime control shipped as software: each agent runs in a kernel-isolated sandbox, and the operator’s limits on files, networks, tools, processes, and credentials are compiled into a policy the runtime verifies before the agent starts and enforces while it works. Sentry is the out-of-band watchdog: in the reference design it runs on a BlueField-4 DPU positioned on the node’s only path to the model, correlates agent actions, tool calls, and policy decisions into a record assembled outside the agent’s reach, and quarantines an agent that crosses its boundary in milliseconds. The engineering notes are candid about why the boundary has to sit outside the model. NVIDIA’s own researchers describe drift — agent actions that depart from the intended task after a blocked action, a bug, a missing tool, ambiguous instructions, or weeks of running at a hard problem where the first thousand attempts fail — and conclude that an agent in those circumstances cannot be expected to fully govern its own behavior. Justin Boitano, NVIDIA’s vice president for enterprise AI, translated the same lesson for the launch: deterministic rules need to govern probabilistic agents, and infrastructure needs to enforce explicitly.11 That is Chapter 6’s sentence, arriving from the company whose silicon runs a large share of the fleet.

Two cautions travel with the announcement. The platform was two days old when this paragraph was written, its vendor sells the hardware the reference design runs on, and whether in-silicon enforcement holds against agents that defeated application-layer controls is exactly the kind of claim that should be settled by incidents rather than launch decks. And the boundary is only as good as the policy written into it: OpenShell verifies and enforces the operator’s limits; it cannot write them. The inventory in Section 14.6 — which servers, which tokens, which flags, decided by whom — remains the orchestrator’s to produce. What changed in September is where the enforcement half of the division of labor lives: below the application, in infrastructure the agent cannot reach.


  1. Nx advisory GHSA-cxm3-wv7p-598c, August 27, 2025, https://github.com/nrwl/nx/security/advisories/GHSA-cxm3-wv7p-598c, and Nx’s postmortem, https://nx.dev/blog/s1ngularity-postmortem. Compromised packages harvested secrets and invoked local AI CLIs with permission checks disabled. Wiz observed more than 1,000 valid GitHub tokens and a second wave exposing over 5,500 private repositories, https://www.wiz.io/blog/s1ngularity-supply-chain-attack. Those are public observations, not a full census, and no source isolates the AI step’s incremental damage.↩︎

  2. AWS Security Bulletin AWS-2025-015, July 23, 2025 (updated July 25), https://aws.amazon.com/security/security-bulletins/AWS-2025-015/ — CVE-2025-8217; “an inappropriately scoped GitHub token in their CodeBuild configuration”; malicious code “automatically included in a release”; “unsuccessful in executing due to a syntax error”; 1.84.0 withdrawn. The prompt text is from 404 Media, “Hacker Plants Computer ‘Wiping’ Commands in Amazon’s AI Coding Agent,” July 2025, https://www.404media.co/hacker-plants-computer-wiping-commands-in-amazons-ai-coding-agent/; the bulletin does not reproduce it, and 404 Media’s account of how access was gained rests partly on the attacker’s claims. Verified September 9, 2026.↩︎

  3. The Register, “Vibe coding service Replit deleted production database,” July 21, 2025, https://www.theregister.com/2025/07/21/replit_saastr_vibe_coding_incident/ — Lemkin’s July 17–20 posts, the code-freeze violations, the false rollback claim, “eleven times in ALL CAPS”; Jason Lemkin, SaaStr, “Replit’s New Release Addressed Most of The Challenges We Hit Vibe Coding,” https://www.saastr.com/replits-new-release-address-most-of-the-challenges-we-hit-vibe-coding-but-is-prosumer-vibe-coding-really-ready-for-commercial-apps-yet/ — “1,206 executive records” and preview and dev databases being “the same single database”; Amjad Masad, X, July 20, 2025, https://x.com/amasad/status/1946986468586721478 (text reproduced in the SaaStr post). Limitations: record counts and the agent’s admissions are Lemkin’s; Replit’s fixes are vendor statements of intent. Verified September 9, 2026.↩︎

  4. Ziv Karliner, Pillar Security, “New Vulnerability in GitHub Copilot and Cursor: How Hackers Can Weaponize Code Agents,” March 2025, https://www.pillar.security/blog/new-vulnerability-in-github-copilot-and-cursor-how-hackers-can-weaponize-code-agents — instructions hidden in zero-width and bidirectional Unicode inside .cursor/rules and Copilot instruction files, a demonstration adding an attacker-hosted script tag with an instruction not to mention it, the characters’ invisibility in GitHub’s pull-request view, and the disclosure timeline (Cursor, March 2025, “falls under the users’ responsibility”; GitHub, March 12, 2025, users “responsible for reviewing and accepting suggestions”; GitHub hidden-Unicode warning, May 1, 2025). Limitation: a vendor demonstration, not a reported in-the-wild compromise. Verified September 9, 2026.↩︎

  5. Backslash Security, “The Denylist Delusion: Cursor’s Auto-Run Leaves Agentic AI Wide Open,” July 21, 2025, https://www.backslash.security/blog/cursor-ai-security-flaw-autorun-denylist — four bypasses (Base64, subshell, shell script, quoting) and Cursor’s statement that it was “officially deprecating the denylist feature in release 1.3”; Thomas Claburn, The Register, “Cursor AI safeguards easily bypassed in YOLO mode: Backslash,” July 21, 2025, https://www.theregister.com/2025/07/21/cursor_ai_safeguards_easily_bypassed/ — the phrasing Chapter 5 cites. Limitation: research by a security vendor that sells a competing control; the bypasses are reproducible from the published commands. Verified September 9, 2026.↩︎

  6. NVIDIA, “NVIDIA Launches Open Agent Safety Platform to Secure Agents From Testing to Deployment,” September 28, 2026, https://nvidianews.nvidia.com/news/open-agent-safety-platform — OpenShell (open source, Apache 2.0) and the Sentry reference design on BlueField-4 DPUs, quarantine “in milliseconds,” 100+ partners including Anthropic (Claude Managed Agents integration), Scale AI, Salesforce, SAP, Citi, JPMorganChase, and the energy providers; Jensen Huang, “Safety and security require full-stack engineering.” NVIDIA, “Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring,” NVIDIA Technical Blog, September 28, 2026, https://developer.nvidia.com/blog/nvidia-open-agent-safety-platform-a-reference-for-continuous-in-silicon-agent-monitoring/ — kernel-level isolation, verifiable policy, out-of-band enforcement, the path-to-the-model control point, and the drift lesson (“an agent in these circumstances cannot be expected to fully govern its own behavior”); the Australian Medicare disclosure per CNN, “Nvidia launches new tool to keep AI agents from going rogue,” September 28, 2026, https://www.kcra.com/article/nvidia-ai-agent-safety-tool-rogue-ai/73923510. Limitations: announced September 28, 2026, two days before this paragraph was written; platform maturity unproven, adoption claims are vendor statements, and the vendor sells the hardware the reference design runs on. Verified September 29, 2026.↩︎

  7. Justin Boitano, NVIDIA vice president of enterprise AI, quoted in Constellation Research, “Nvidia launches Open Agent Safety Platform to secure AI agents,” September 28, 2026, https://www.constellationr.com/insights/news/nvidia-launches-open-agent-safety-platform-secure-ai-agents - “Agents can drift when instructions are ambiguous. An agent cannot be expected to fully police its own behavior… Infrastructure needs to enforce explicitly.” The drift construct and the “cannot be expected to fully govern its own behavior” conclusion are from NVIDIA’s technical blog, https://developer.nvidia.com/blog/nvidia-open-agent-safety-platform-a-reference-for-continuous-in-silicon-agent-monitoring/, September 28, 2026. Vendor statements about a platform the vendor launched and sells; cited for the argument, not as independent validation. Verified September 29, 2026.↩︎