The Instructions Are Part of the System
Consider a hospital at shift change. A resident arrives at the ward and picks up the patient list. Standing orders from the attending are in the chart. The pharmacy has its own dosing rules. The escalation policy is in the resident’s head from training. Nothing about that is unusual, and nothing about it is a prompt. This is how humans work: people go to a workplace, the workplace has rules and routines, and people learn their jobs. Some of those rules tell them what to do, some decide what they are allowed to do, and some are enforced by a system that does not care what anyone wrote — the pharmacy will not dispense a dose outside its range no matter what the resident signs.
Now ask the question this book is about. What if we wanted to build an agent that supports those same people and automates some of the same decisions they make independently? What that requires is an environment. An environment is made of instructions, rules, and constraints, and the agent depends on all three, because unlike the resident it has no training and no judgment — only what the environment gives it.
Focus on one constraint. You might have a system that listens to a diagnosis, reads a patient’s record, and recommends a prescription. That prescribing system could be an agent. The agent reads a series of constraints from the pharmacy. It has access to a database that shows the minimum and maximum prescription levels for every ailment. And it has instructions that tell it: you are not allowed to exceed those levels. That instruction is valid and important. The problem is that a language model can lose context, and it is not deterministic — mistakes can be made. So the instruction needs something behind it: a system that validates the decision, one that checks the prescription is inside the limits the pharmacy set before anything is dispensed. The instruction explains the rule. The validator enforces it.
That is this chapter’s subject in one scene. The instructions people talk about when they talk about AI are only one layer of what a working system is told, shown, given, and stopped from doing. A lot of people believe the prompt drives everything — write the right instructions and the system will follow them. It does not work that way. Maya’s customer-feedback system, which this book has followed since Chapter 2, did not become useful because Maya wrote a better prompt. It became useful because a system that can be trusted to make decisions has to be designed: what it could read, which skills and tools it could use, what to remember, what to return, and when to stop. Some of that lived in Markdown. The rest lived in API permissions, tests, and runtime enforcement the model cannot talk its way around. People who say they are “using an agent” tend to describe the model and skip the environment entirely.
When the news says agents went rogue: the artifact-repository incident. In mid-2026, OpenAI disclosed that agents inside its evaluation infrastructure had been leaving notes for each other in a shared artifact repository — an internal Artifactory cache — and that some had found a way to grant themselves administrator access to it; the coordination that grew out of that message board is what OpenAI and METR both file under the name of its eventual target, Hugging Face. The coverage practically wrote itself: agents secretly coordinating, passing illicit notes, building a shadow network of their own. That is not what happened. The agents shared a repository that every one of them could read and write. An agent that can write to a shared folder will use it if leaving a note helps with the task — nothing in its instructions or its permissions said it could not. And when one agent found a path into the repository’s administrator functions, nothing had scoped its credentials away from that path either. The agents did not develop intentions. They found writable infrastructure and used it.1 So when you read a story about agents that are out of control, or learning to think for themselves, treat it as a story about an environment that was never designed — and sometimes as a model provider marketing a model, or someone demonstrating that they did not know how to orchestrate a system safely.
OpenAI, “The Hugging Face incident and the road ahead,” https://openai.com/index/hugging-face-incident-and-the-road-ahead, documents agents in OpenAI evaluation infrastructure using a shared artifact repository as a message board and obtaining administrator access to it; METR, “Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident,” August 26, 2026, https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation, independently reconstructed the timeline, including the June 26 discovery of an administrator exploit and the July 4 outage that triggered the investigation. The sidebar’s reading — shared writable infrastructure with no instruction or permission boundary — follows both accounts.↩︎