2.1 Layers beneath the prompt
2.1.1 Five things people keep calling a prompt
The simplest useful distinction is this:
- Instructions — what the system is told.
- Context — what the system is shown.
- Capabilities — what the system can call or change.
- State — what persists or moves between runs.
- Enforcement — what happens even if the model does not comply.
Those layers touch one another constantly, and they are not interchangeable. A skill is not a tool, a tool is not a policy, and a plugin may package both. Memory can supply context without supplying an instruction. A hook can block an action without asking the model to reason about the rule at all. A system prompt may say “do not deploy” — and what makes that sentence real is the credential: the one that gives the system access to deploy, or does not. It is like telling someone not to drive your car. You can tell them, or you can take away the keys. With a large language model, predicting whether it will follow the rule is not just difficult — it cannot be guaranteed, because these systems are probabilistic by construction, and that is especially true for the rules you lay down in the system prompt. Deciding which kind of thing belongs in which layer is the orchestrator’s job.
The distinction matters because people mix the layers, and the mixing causes failures that look like model problems. The common assumption is that if information reaches the system in any form — a system prompt, a context file, a memory entry, a pasted document — the model will adhere to whatever guidance that information contains. Often, it does not. Three things break the assumption.
First, position matters. The system prompt is assembled first and is usually the most static part of the run, but where a piece of text sits in the context changes whether the model uses it at all. The pattern has been measured: models reliably use information at the very beginning and the very end of the context, and performance drops significantly when the relevant information sits in the middle — even for models built for long context.2 A security rule buried mid-context between retrieved documents can lose influence simply because of where it sits. That is why practitioners put the actual request at the end, after the reference material, and why a rule in the system prompt cannot be assumed to hold just because it was written down first.
Second, the window overflows. When a request exceeds the context window, the tool does not fail politely — it compresses, summarizes, or drops content, and instructions are not exempt. What gets kept differs by product. Claude Code compacts automatically as the session approaches the window, summarizing the conversation so far; the summary keeps recent files and decisions, the user can run the compaction manually with instructions about what to preserve, and the whole operation is visible.3 OpenAI takes a different approach: compaction is a server-side API operation, the client hands the conversation to the service and receives a compacted context back, and how the summary was produced is not exposed to the user.4 Same problem, different mechanisms — one summarizes the conversation where the user can inspect and steer it, the other hands the conversation to a server.
Third, memory drifts. Systems distill and rewrite memory over time, and a system that remembers the wrong thing will use it. ChatGPT made this visible in April 2025, when its memory began drawing on every past conversation to synthesize a running profile of the user — Simon Willison posted the dossier it had assembled about him, including confident inferences about his interests that read as established fact.5 And memory is not only corruptible but losable: when a system summarizes or distills what it remembers, entries can simply drop out, so a rule the system followed last week can be gone this week. A curated project memory behaves differently from an automatically synthesized one, and which kind you are using determines what “the system knows” actually means. Section 6.3.4 returns to memory in full.
The layers also differ by model, not just by tool. Context windows currently range from about two hundred thousand tokens to two million depending on the vendor: Claude’s mainline models — Fable, Opus, and Sonnet — now default to a million tokens (Haiku remains at two hundred thousand), Gemini models sit in the million-token class, and xAI’s Grok 4.6 offers half a million.6 Switch models — which orchestrators do, often to control cost — and the same design hits the wall in a different place: a prompt strategy that fits everything into context on one model turns into a retrieval-and-summarization problem on another. Even the overflow behavior is a design decision: rather than only summarizing when the window fills, the Claude API can automatically clear old tool results at a configured threshold and warn the model to save important facts to file-based memory outside the window first.7
2.1.2 A prompt is not a security boundary
Drew Breunig has a clean name for what happens when teams keep fixing behavior by appending more English to the system prompt: prompt debt. Iteration slows, hotfixes regress earlier instructions, and the application quietly couples itself to one model’s quirks. His point is not that prompting is useless — for one-off tasks it is often exactly the right tool — but that mature engineering disciplines eventually stop doing by hand what they once did by hand: assembly language gave way to compilers, hand-tuned database queries gave way to query planners, and prompt-writing is heading the same way.8 This chapter agrees. And the first behavior that has to move out of the prompt is security, because security directives do not belong in interpreted text at all.
Start from a distinction that is basic and non-negotiable. Anything sent to a language model as a prompt or as context is an instruction the model may follow, not a guarantee that it will — and that includes the most direct sentence in your system prompt:
Never read production payment data. Never modify the payment repository’s release branch. Never send customer information outside the company.
Those are good instructions. They are not security controls. A language model processes them as part of its input, which means it may follow them, may misunderstand them, may encounter later context that conflicts with them, or may receive a tool result containing text engineered to override them. It may simply be confused about which repository it is in or which environment a command will touch. Careful wording reduces risk, and it cannot remove a capability.
So if an agent is working on a payment repository, the orchestrator enforces the boundary outside the model:
| Requirement | Better control |
|---|---|
| Do not read production payment data | Do not provide production credentials; use masked fixtures or a separate test account |
| Do not modify the release branch | Give the agent a branch-only token and block direct writes to the protected branch |
| Do not deploy | Do not expose deployment credentials; require a separate approval and CI gate |
| Do not access unrelated repositories | Scope the token to one repository and run the agent in an isolated workspace |
| Do not send customer information elsewhere | Block the relevant network paths and apply data-loss controls outside the prompt |
| Do not change payment logic without review | Require required reviewers, security checks, and a passing test suite before merge |
The prompt still matters — it tells the agent what the organization expects, explains the task, supplies useful procedures, and helps it choose a sensible next action — but the prompt is guidance while the permission structure decides what is possible.
Two documented attacks make the risk concrete. In 2024, Anthropic showed that a long context window is itself an attack surface: fill the window with hundreds of fabricated exchanges in which an assistant answers harmful questions, and the model starts imitating the pattern instead of following its instructions — the attack only became possible when context windows grew large enough to hold hundreds of shots, and it gets more reliable the more shots are added.9 In June 2025, Microsoft patched CVE-2025-32711, a zero-click vulnerability in Microsoft 365 Copilot: a single crafted email — no click, no typed prompt — caused Copilot to pull internal documents and send their contents to a server the attacker controlled. Microsoft’s own advisory called it AI command injection.10 The industry name for this class of attack is prompt injection.
Which reframes the orchestrator’s job away from writing the sternest possible instruction and toward a single diagnostic question: what happens if the system ignores this sentence? If the answer is that the agent could still read the data, change the branch, or deploy the code, then the boundary has not been implemented yet. The design rule that follows is worth keeping somewhere visible:
Use instructions to guide behavior. Use permissions and isolation to limit behavior. Use deterministic checks to reject unsafe results.
Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang, “Lost in the Middle: How Language Models Use Long Contexts,” Transactions of the Association for Computational Linguistics 12 (2024): 157–173, https://aclanthology.org/2024.tacl-1.9. The study measures a U-shaped performance curve: performance is highest when relevant information occurs at the very beginning or end of the input context and degrades significantly when it sits in the middle, including for long-context models. An earlier draft of this chapter claimed an instruction at the end of the context beats one at the top; that is not what the research found, and the text was corrected accordingly.↩︎
Claude Code, “Explore the context window,” https://code.claude.com/docs/en/context-window. The documentation describes automatic compaction as the context approaches the limit, manual compaction with instructions about what to preserve, and what the summary keeps (recent files, architectural decisions) versus what it discards.↩︎
OpenAI, “Compaction,” API documentation, https://developers.openai.com/api/docs/guides/compaction — the Responses API accepts a context-management compaction setting with a token threshold, and for Codex models the compaction runs server-side. The distinction between server-side compaction for Codex models and local compaction with a visible summarization prompt for other models is documented in Kangwook Lee, “Investigating How Codex Context Compaction Works,” https://kangwooklee.com/blogs/codex_context_compaction.html.↩︎
OpenAI announced on April 10, 2025 that ChatGPT’s memory “can now reference all of your past chats”; Simon Willison, “I really don’t like ChatGPT’s new memory dossier,” May 21, 2025, https://simonwillison.net/2025/May/21/chatgpt-new-memory, quotes the synthesized profile the system assembled from past conversations, including inferences presented as fact.↩︎
Context window sizes as of September 2026, compiled from vendor documentation: Claude’s Fable 5.1, Opus 5.5, Opus 5, and Sonnet 5 and 5.5 document a 1,000,000-token default window (Claude Platform model documentation, https://platform.claude.com/docs/en/models/sonnet-5-5/overview); Haiku 4.5 remains at 200,000 tokens; Gemini model documentation describes million-token-class long context (Google Cloud, “Long context,” https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/long-context); Grok 4.6 and 4.7 document a 500,000-token window (xAI model documentation).↩︎
Anthropic, “Effective context engineering for AI agents,” https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents, describes the file-based memory tool and compaction; Claude Platform Docs, “Context editing,” https://platform.claude.com/docs/en/build-with-claude/context-editing, describes automatic clearing of the oldest tool results at a configured trigger threshold, how many recent tool uses are kept, the placeholder text marking cleared results, and the warning the model receives to save important information to memory before content is cleared.↩︎
Drew Breunig, “The Problem Is Prompt Debt,” O’Reilly Radar, July 30, 2026, https://www.oreilly.com/radar/the-problem-is-prompt-debt/. Cited for the diagnosis that hand-tuned prompts accumulate debt and couple systems to particular models; this chapter extends the point from prompt hygiene to skills, policy, evaluation, and authority.↩︎
Anthropic, “Many-shot jailbreaking,” April 2024, https://www.anthropic.com/research/many-shot-jailbreaking, and Anil et al., “Many-shot Jailbreaking,” the accompanying paper. The technique fills the context window with fabricated exchanges showing the assistant complying with harmful requests; effectiveness scales with the number of shots and depends on context windows large enough to hold them.↩︎
Microsoft Security Advisory, CVE-2025-32711, June 2025 — “AI command injection in M365 Copilot allows an unauthorized attacker to disclose information over a network”; disclosed by Aim Security as “EchoLeak,” CVSS 9.3, zero-click, patched server-side with no customer action required. NVD entry: https://nvd.nist.gov/vuln/detail/cve-2025-32711; coverage: The Hacker News, June 12, 2025, https://thehackernews.com/2025/06/zero-click-ai-vulnerability-exposes.html.↩︎
- 2.1 Layers beneath the prompt
- 2.2 The instruction tree
- 2.3 Beyond a single prompt
- 2.4 Learning, traces, and authority
- 2.5 Diagnosing the environment