4.1 Variation and fences

Book 2 · The Delegation ContractChapter 4 · section 1 of 5

4.1.1 The new uncertainty

The safety machinery software engineers rely on — tests, logs, alerts, rollbacks, incident procedures — rests on one assumption: given the same inputs and the same environment, the code produces the same result. Losing that assumption costs more than novelty. Tests assume a failure can be reproduced by running it again; logs assume the path through the code explains the output; alerts assume a threshold crossed once means something. None of those assumptions survives a component that can return a different answer to the same input — not because the machinery is wrong, but because it was built to explain deterministic behavior. The machinery for nondeterministic behavior has to be designed around variation specifically: classify it, contain it, and record enough to defend it.

Orchestration adds a subtler failure on top. A system can pass the test somebody wrote while failing the requirement nobody remembered to write down, because the test pinned one behavior and the model produced another, equally plausible one.

The useful distinction is not whether the code is deterministic or nondeterministic but whether a given variation changes meaning, because treating all nondeterminism as hallucination is far too coarse to operate on. An orchestrator has to classify the variation before deciding what the system is permitted to do next, and there are at least four categories worth naming.

  • Cosmetic variation changes presentation but not meaning: a different word for the same apology, a different ordering of the same facts.
  • Interpretive variation changes emphasis and may require review: a medical summary that calls a finding “reassuring” instead of “stable.”
  • Factual variation changes what the system claims to know: a wrong date, a wrong recipient, a wrong regulation.
  • Operational variation changes what the system does: a denied claim, a closed case, a rerouted shipment.

A reader will reasonably ask whether those four categories are established or invented, and the honest answer is that the words are ordinary and the arrangement is this book’s. The research literature classifies wrongness extensively. The canonical division separates hallucinations that contradict the source from those that cannot be verified against it, and the most-cited survey for language models reorganizes that into factuality versus faithfulness, with faithfulness split further into instruction inconsistency, context inconsistency, and logical inconsistency.5 Those taxonomies sort what is wrong with a claim. The four categories above sort a different axis: what changes about the world when the output changes, from wording nobody will notice to a decision a person depends on. Ordering variation by consequence is what makes the classification operational — it tells the orchestrator which fence the variation has to hit.

A side note on why this book does not just call them guardrails. A guardrail is what an alpine slide has: a steel rail along a fixed track, built for a sled with no steering and no judgment, expected to bang into the rail every run forever. That is an accurate picture for evaluation loops that catch a system after it has already gone wrong, and the wrong picture for an orchestrated system. Your agents are not coaster cars; they are entities expected to learn the boundary — routing around a constraint, asking before acting, flagging an input the fence rejects. A system still slamming into its guardrails every week is not being orchestrated; it is being caught. The word fence is deliberate: a fence gives a herd a boundary it can live inside and learn, and a herd constantly hitting it means the shepherding is failing, not that the fence is working.6

The honest version of this comparison gives guardrails their due: where you want dumb, tireless rejection — a validator on a structured field, a schema check — the rail is exactly right, and you should keep it. The distinction that matters is not hard controls versus soft ones; it is whether the control is the only thing between the system and the harm, or whether the system is expected to internalize the boundary over time. Audit every control in a topology for which kind it is.

Controls should tighten as variation moves from cosmetic toward operational, which is more than a confidence score can tell you on its own. The system needs to know what kind of error it is looking at and what the consequence would be if that error escaped, and that is the whole distance between “the model is unsure” and “the evidence conflicts across two sources, so the workflow pauses and requests a domain review.”

4.1.2 Fences around an unpredictable output

A fence sets limits on what the system is allowed to do when the available evidence is incomplete, conflicting, or unexpected — whether it can proceed, must take a safer action, or needs to involve a human. The three fences below are a practical starting set, not the only ones. Domains add their own: a clinical workflow may need a consent fence; a payments system may need a settlement-window fence; a public-sector system may need a statutory-authority fence. What matters is that each fence answers a question the model cannot answer for itself.

Acceptable outcome range. What must remain invariant, and what is allowed to be creative? A customer-service system may vary its wording and may not invent a return policy; a research workflow may explore several hypotheses and may not quietly promote a guess to a fact. The range is set by the domain rather than by the model’s confidence, which is why varying the rhythm of a translation is harmless while varying the recipient of a medical shipment is catastrophic. “Sometimes our systems send the wrong prescription medications” is not an acceptable class of hallucination.

Evaluation. An evaluation set should cover adversarial, routine, and badly formed inputs, including the requests people submit when they are tired, angry, unclear, or deliberately pushing the system outside its purpose — because whatever cases the set contains are the behaviors that will receive attention. Evaluate a customer-service workflow only for politeness and you have learned nothing about its accuracy, its privacy protections, its escalation decisions, or its ability to solve the customer’s problem.

So build the set from real operating conditions: common cases, boundary cases, past incidents, adversarial attempts, and situations where refusal is the correct answer. Hold back some cases the system never saw during development, so the evaluation measures general performance rather than familiarity with known tests. Then review the set on a schedule, checking which scenarios are missing, where the examples came from, and whether the people affected by the system are adequately represented, because a system can pass every defined test and still perform badly for the users and situations the evaluation forgot.

Intervention path. An intervention mechanism is worth having only if somebody has both the authority and the practical ability to use it, under pressure, without stopping to read documentation. The process should pause the workflow before an irreversible action, present the evidence that triggered the pause, show the available alternatives, record who decided and why, and allow the action to be reversed where reversal is possible. Rehearse each of those steps regularly, too. Until operators have actually practiced stopping the system, nobody can honestly claim to know whether the intervention path works.

Fences also have a history, because the idea of a computer holding a boundary around an unpredictable operator did not start with language models. The Airbus A320 entered service in 1988 with flight envelope protection built into its fly-by-wire software: the control computers refuse a stick input that would stall the aircraft or overstress it, no matter who is commanding it or how tired they are. Boeing took a different position on the 777 and allowed the crew to override the envelope. Both manufacturers put the boundary in the deterministic layer — in the computers, not in the judgment of whoever was flying — and the argument they left open, hard boundary or soft, is the same argument an orchestrator inherits when deciding whether a model can talk its way out of a constraint. The A320’s boundary held when it mattered most, and it also constrained: the NTSB found that during the 2009 Hudson River ditching, the aircraft’s alpha-protection mode kept the plane inside its aerodynamic limits while the crew worked the forced landing — and that the same protection kept the captain from reaching the maximum angle of attack he was commanding with full aft stick in the flare, so the touchdown was harder than it might otherwise have been.7


  1. Prior art for classifying wrongness: Ziwei Ji, Nayeon Lee, Rita Frieske, et al., “Survey of Hallucination in Natural Language Generation,” ACM Computing Surveys 2023 (intrinsic versus extrinsic hallucination); Lei Huang, Weijiang Yu, Weitao Ma, et al., “A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions,” ACM Transactions on Information Systems 2024, https://arxiv.org/abs/2311.05232 — “We categorize hallucination into two primary types: factuality hallucination and faithfulness hallucination,” with faithfulness subdivided into instruction, context, and logical inconsistency. The four-category arrangement in the text is this book’s.↩︎

  2. “Fence” as a label for these controls is this book’s; the industry’s word is guardrails, used both by the product category (NVIDIA’s NeMo Guardrails, cited in Section 8.2.1) and by the research surveys of the field. The livestock metaphor is deliberate: a fence contains a herd without predicting any animal in it, which is the working relationship between an orchestrator and a nondeterministic component.↩︎

  3. Airbus, “Safety Innovation #7: Flight Envelope Protection,” February 2023, https://www.airbus.com/en/newsroom/stories/2023-02-safety-innovation-7-flight-envelope-protection, and Wikipedia, “Flight envelope protection,” https://en.wikipedia.org/wiki/Flight_envelope_protection. The A320 (1988) was the first commercial aircraft with full envelope protection; Boeing’s 777 lets the crew override the envelope. On US Airways Flight 1549 (2009), the NTSB found the A320’s alpha-protection mode was active in the last 150 feet of the descent and that, because of it, “the airplane could not reach the maximum angle of attack” attainable in normal law despite the captain’s full aft sidestick, while still providing maximum performance for the weight and configuration; the report also notes the protections let the captain pull full aft without risk of stalling. NTSB, Aircraft Accident Report AAR-10/03 (May 4, 2010), https://www.ntsb.gov/investigations/AccidentReports/Reports/AAR1003.pdf.↩︎