5.5 What security cannot do

Book 2 · The Delegation ContractChapter 5 · section 6 of 7

An honest chapter says what the controls do not cover. You cannot make the model itself incorruptible. A language model processes instructions and data in the same channel, and there is no configuration that makes injected text unable to influence it. Every control above works by accepting that fact and shrinking what influence can reach: not the model, but what the model can touch. The instructions chapter stated the design rule; the security addendum here is that the model’s fallibility is a constant of the design, not a bug to fix.

That reframes the security question the orchestrator answers: not “can this system be tricked?” (yes, accept it) but “what is the worst outcome if it is?” If the answer is acceptable — contained by the stack above — the system is deployable. If the answer is exfiltration of customer data or an unauthorized payment, the boundary has not been implemented yet.

This is also where the introduction’s caution — that the governance layers and permission envelopes around these systems are mostly early, rough, or missing — applies with particular force. The controls are known; the stacked discipline around a compromised agent is a design practice the industry is still standardizing. OWASP’s first agent-specific risk list is dated December 2025, and the incidents it cites as examples are the ones in this chapter. U.S. and allied cyber agencies have published joint guidance on deploying AI systems securely — sandboxing, credential management, zero-trust architecture, adversarial testing — built on existing frameworks rather than a new AI-specific regime.14 The security practices for agent systems are the security practices for privileged software, applied to a component that reads text.


  1. NSA Artificial Intelligence Security Center, CISA, FBI, ASD ACSC, CCCS, NCSC-NZ, and NCSC-UK, “Deploying AI Systems Securely” (joint Cybersecurity Information Sheet, April 2024), https://www.cisa.gov/news-events/alerts/2024/04/15/joint-guidance-deploying-ai-systems-securely — sandboxing, credential management, zero-trust architecture, adversarial testing. Verified September 9, 2026.↩︎