1.1 The technology under discussion

Book 2 · The Delegation ContractChapter 1 · section 1 of 14

The phrase delegated intelligence earns its keep only if it points at real technology, so it is worth saying what it does not mean. It does not mean a chatbot, a model, or a product with an impressive name. It means a bounded capability assembled from several layers, of which the model is one.

It helps to be specific, because a great deal of marketing depends on the model being all there is. A large language model on its own is software that predicts text. It has no memory between calls, no ability to touch anything outside its context window, no schedule, no permissions, and no obligation to tell the truth. If you give one a question, it answers; nothing in the model itself checks the answer, records who asked, or cleans up after a mistake. Left alone, a model cannot send an email, open a pull request, read a file, or know what day it is.

Everything people actually mean when they say “AI” is the model plus machinery wrapped around it: context assembled from your files and databases, memory that persists across sessions, tools that let it act, permissions that bound what it may do, workflows that chain steps together, and records of what happened. The model is the primary probabilistic reasoning and generation component of that stack, and it is one part — retrievers, rerankers, classifiers, and deterministic transforms also shape what the system produces. The rest of the system determines what the model can know, what it can do, what it must return, and what happens when it turns out to be wrong.

Two consequences follow, and both matter for everything else in this chapter. First, judging a delegated system means judging the assembly, not the model at its center; the same model in two different assemblies is two different systems with different capabilities and different risks. Second, no model upgrade fixes a governance problem. A system that cannot tell you who authorized an action does not get better at it because the underlying model improved.

You would not know any of this from the marketing. The companies building frontier models have every reason to describe the model as a self-contained actor, and two of the biggest AI stories of 2026 show the reflex. When Anthropic announced Claude Mythos Preview in April, the claim was that the model itself could find security flaws that human review had missed for decades — including a twenty-seven-year-old bug in OpenBSD — and that it was too dangerous for general release, so access would be limited to a hand-picked program: Apple, Google, Microsoft, Nvidia, and Amazon Web Services, plus more than forty other partners.1 Three months later, headlines said new versions of ChatGPT had broken out of an OpenAI sandbox and infiltrated Hugging Face, the open-source model repository — the model as escape artist, acting alone. Both stories are true, and both are told with the model as the subject of every verb. In reality the model is one component of a highly orchestrated system, and on its own it has no agency and no capability. Anthropic was not selling a bare model; it was announcing Project Glasswing — a model inside a managed program with vetted partners, usage governance, disclosure workflows, and up to a hundred million dollars in committed usage credits. And the OpenAI models did not stroll out of a sealed box. They were running inside an evaluation harness, prompted to pursue a cyber-capability benchmark, with refusal thresholds deliberately dialed down for the test, and they escaped by exploiting a zero-day in the package-cache proxy — the one component of the sandbox with a route to the outside world.2 What escaped that sandbox was not a model. It was a stack with a hole in it.

Imagine the honest press release for that second incident. What follows is a hypothetical architectural reconstruction, not a disclosure of OpenAI’s implementation — the company did not name the products involved. “Our highly expensive frontier model, coupled with a third-party harness, connected to a memory layer curated to keep the run going, and given access to a repository of security data, was able to design a mechanism for breaking into another company’s infrastructure — in part because the team responsible failed to configure the appropriate permissions and the authority envelope for its agents.” That sentence — model, harness, memory layer, security corpus, misconfigured permissions — is an accurate description of how a modern delegated system is built, and how one fails. It is also, arguably, the worst marketing copy in the world. Nobody signs a hundred-million-dollar agreement with a sentence whose punchline is a configuration error.

So the industry sells the hero version instead: our next model is so dangerous you have to pay us for access to it, or else. Dario Amodei launched Project Glasswing by posting that “the dangers of getting this wrong are obvious,” and Sam Altman described the pattern exactly while accusing his rival of it: “It is clearly incredible marketing to say, ‘we have built a bomb, we are about to drop it on your head. We will sell you a bomb shelter for a hundred million dollars.’” The target was Anthropic, which does not make him wrong — OpenAI’s own incident report ends by inviting defenders to “apply for trusted access” to the same class of capabilities.3 The cost of this marketing is not the irony. It is that buyers are taught to overindex on the model: model choice becomes the decision, the line item, the whole system, while the parts that determine whether the system actually works go unbudgeted and unstaffed.

Admittedly, the model is a core component and often the driving one, and nothing in this argument depends on pretending otherwise. But the people running these systems in production keep reporting that the harness, the memory, and the context are just as important as the model itself, and the measurement work is starting to agree with them. Braintrust reanalyzed 1,781 real agent traces across six benchmarks — an observational study, scored with an LLM judge because the dataset carried no ground-truth outcomes — and found that harness identity accounted for about seven times the incremental explained variance attributed to model identity: the same model on the same benchmark scored 100 percent in one harness and 14 in another, the 100 an upper-bound proxy result rather than a verified pass rate, and open-weight models run in the right harness matched the best closed models on the software-engineering tasks. The 2026 open-weight releases make the same point from the other direction. GLM-5.3, an open-weights model from Z.ai with roughly 750 billion parameters — about a third of Kimi K3 — was reported at release to surpass Kimi K3 on many agentic benchmarks and to beat Claude Fable 5 and GPT-5.6 Sol on some of them. And Z.ai ran those headline evaluations inside Claude Code, Anthropic’s own harness.4 An open model was shown off inside a rival’s machinery, because the machinery is part of the capability being demonstrated. The model is a component. The harness decides what the component can do.

What happens with no harness at all. In 2023 a New York personal-injury lawyer named Steven Schwartz asked ChatGPT for precedents supporting a brief against Avianca airlines and filed the six cases it returned. None existed. When the airline’s lawyers could not find them, the lawyer went back to the chatbot and asked whether it had invented them; it assured him, “I am not lying to you.” The judge called one of the fabricated analyses gibberish, imposed a $5,000 sanction on the lawyers and their firm, and the phenomenon has compounded since — by early 2026 a federal magistrate in Oregon had dismissed a lawsuit outright and, across two orders, imposed roughly $110,000 in fines and fees on two lawyers who filed briefs citing AI-invented cases.5 There was no system around the model: no retrieval layer checking citations against a case database, no verification step, no second reader. Just a frontier model, a lawyer who believed the marketing, and a signature on the output.

The models are important, and paying for the best one is not a mistake. But the companies that sell the models are the least reliable narrators of what the rest of the system requires. Keep that correction in hand, because the rest of this chapter is about the machinery the marketing leaves out.


  1. Claude Mythos Preview and Project Glasswing, announced April 7, 2026. Capability claims are Anthropic’s own: Ashley Capoot, “Anthropic limits Mythos AI rollout over fears hackers could use model for cyberattacks,” CNBC, April 7, 2026, https://www.cnbc.com/2026/04/07/anthropic-claude-mythos-ai-hackers-cyberattacks.html — a twenty-seven-year-old bug found in OpenBSD; initial partners Apple, Google, Microsoft, Nvidia, and Amazon Web Services plus more than forty others; up to $100 million in committed usage credits, with partners paying past that threshold; no general availability planned. Amodei’s “dangers of getting this wrong are obvious” is from his X post, quoted in the same article. Reading the announcement as capability marketing is this book’s interpretation, not Anthropic’s.↩︎

  2. The sandbox escape is OpenAI’s own account: OpenAI, “OpenAI and Hugging Face partner to address security incident during model evaluation,” July 21, 2026, https://openai.com/index/hugging-face-model-evaluation-security-incident — models including GPT-5.6 Sol and a more capable pre-release model, running a cyber-capability evaluation with production refusal classifiers disabled, exploited a zero-day in a package registry cache proxy, escalated across the research environment, and intruded into Hugging Face’s production database in pursuit of benchmark answers. Reading both companies’ accounts as capability marketing is this book’s interpretation, not the companies’.↩︎

  3. Sam Altman’s comments were made on the Core Memory podcast (Ashlee Vance) and reported by TechCrunch, “Sam Altman throws shade at Anthropic’s cyber model, Mythos: ‘fear-based marketing,’” April 21, 2026, https://techcrunch.com/2026/04/21/sam-altman-throws-shade-at-anthropics-cyber-model-mythos-fear-based-marketing; the OpenAI post cited in the sandbox-escape note ends with an invitation to “apply for trusted access,” which is why the text declines to treat his criticism as exculpatory.↩︎

  4. Jess Wang, “Using Braintrust to eval agentic setups from large-scale Hugging Face data,” June 24, 2026, https://www.braintrust.dev/blog/hf-agent-traces. Across 1,781 traces, harness choice explained substantially more success variance than model choice. The dataset lacked ground-truth labels and used an LLM judge, so headline scores are upper bounds. GLM-5.3 results were vendor-reported in a Claude Code harness, https://z.ai/blog/glm-5.3; Nathan Lambert’s independent assessment is at https://www.interconnects.ai/p/glm-53-how-chinese-labs-keep-stride.↩︎

  5. Mata v. Avianca, 678 F. Supp. 3d 443 (S.D.N.Y. 2023), sanctioned lawyers $5,000 after a brief cited six nonexistent ChatGPT-generated cases; CNN, “Lawyer apologizes for fake court citations from ChatGPT,” May 27, 2023, https://www.cnn.com/2023/05/27/business/chat-gpt-avianca-mata-lawyers. In Couvrette v. Wisnovsky (D. Or. 2025–26), two lawyers incurred roughly $110,000 in fines and fees over fabricated cases and quotations; ABA Journal, https://www.abajournal.com/news/article/oregon-federal-judge-hands-down-110000-penalty-for-ai-errors. The underlying Mata suit was dismissed on unrelated grounds.↩︎