4.4 Training the envelope
4.4.1 The claims workflow: defining and training the rules
The Castellano test from earlier in this chapter showed the problem in miniature: three separate agents, three different operational answers, all of them plausible — unmanaged variation when a consequential decision sits on a nondeterministic path.
For the claims workflow, the rules began as a short failure envelope:
| Question | Rule the team wrote first |
|---|---|
| What may vary? | Wording of document requests; which supporting note is emphasized |
| What must never vary? | No denial without a named human; no fraud inference from a protected characteristic; no case closed by the model alone |
| What signals drift? | Average accuracy and confidence calibration |
| What requires a second opinion? | Confidence below a floor; conflicting evidence |
| What action is reversible? | Sort, draft, request missing documents, recommend a queue |
| What action is not? | Denial, fraud flag, case closure |
Writing the envelope was the first pass. The second was training it: an evaluation set of ordinary claims, incomplete claims, unusual-but-legitimate claims, and a handful of past cases where the old process had made a costly mistake. On the dashboard the system looked fine. Throughput was up. Satisfaction on the routine path was up. Average accuracy was fine.
Then a regional audit asked who was being sent to manual review, and why. Claims from one region — served by a small set of local payroll and clinic providers — were overrepresented in the queue. The model had learned to treat unfamiliar documentation patterns as suspicious, because unfamiliarity had correlated with errors in the training data. The average metric had no category for that. The rules for making decisions were present; the rules for detecting which variation mattered were incomplete.
That failure mode is not hypothetical, and it does not require a language model. In 2019 a team led by Ziad Obermeyer audited a commercial algorithm applied to roughly 200 million people a year in the United States, used to decide which patients should be referred into extra care-management programs. The algorithm used predicted healthcare cost as a stand-in for health need. Because less money had historically been spent on Black patients at the same level of illness, the algorithm concluded they were healthier, and the bias reduced the number of Black patients identified for extra care by more than half. Every average metric looked fine — exactly the way the claims dashboard looked fine — because the metric had no category for the harm it was doing.25 The drift rule that catches this class of failure has to be designed in deliberately, from the operating record, after asking who the average leaves out.
So the orchestrator stopped the affected route, reviewed the queued cases against the original evidence, rebuilt the evaluation set with a regional stratum, added a regional reviewer, and rewrote the envelope: drift now included routing mix by region, provider, and documentation pattern; second opinion now included an unfamiliar documentation signature; and any case below that floor had to escalate to a named adjuster. The model could still sort and request, but it could no longer treat unfamiliarity as suspicion without a human. The system had less autonomy afterward, and more trust, because the organization had learned what the previous autonomy had been hiding.
4.4.2 Human factors: why the control breaks under load
A fence that requires perfect attention at the exact moment of failure is a weak fence. The person at the console needs enough context to understand what is happening, enough authority to act, and enough time to decide — and a system that presents fifty alerts, each marked urgent, has not preserved human oversight so much as outsourced the filtering to panic.
A 2026 synthesis by Vahid Garousi names the resulting condition cognitive overload: the ongoing need to review, validate, repair, and integrate AI-generated artifacts, combined with the volume of suggestions, alternatives, prompts, and outputs that any one reviewer can absorb.26 The paper identifies decision density, context switching, and the sustained evaluation of superficially plausible alternatives as the mechanisms — ordinary working conditions for anyone whose tools produce more than they can verify.
The human-factors literature from aviation describes the same condition under a different name. Automation-induced complacency, defined in a 1993 study by Singh, Molloy, and Parasuraman, is “self-satisfaction which may result in non-vigilance based on an unjustified assumption of satisfactory system state.”27 The 2010 Parasuraman and Manzey review showed the effect persists across modern automated systems, and the MITRE 2016 review “Nothing Can Go Wrong” tied it to the 1992 Air Inter A320 crash near Strasbourg, where the crew, with the autopilot still in vertical-speed mode rather than flight-path-angle mode, dialed in what they intended as a 3.3-degree descent and got 3,300 feet per minute instead, into a ridge no ground-proximity warning announced. What travels from that literature is the construct rather than the accident: the more reliable the automation, the harder it is to stay attentive, and the more catastrophic a missed signal becomes.
A 2024–2026 body of human-in-the-loop review work converges on four failure modes that travel well from aviation and medicine into orchestration.28
| Failure mode | What it looks like in an orchestrated system |
|---|---|
| Automation bias | The reviewer accepts the model’s recommendation because it arrived with confidence and structure, and would have rejected the same evidence from a human colleague. |
| Alert fatigue | The reviewer ignores the alert because there are too many of them and most are noise; the genuine signal looks identical to the noise. |
| Scalability limits | The reviewer can only attend to a fraction of the cases the system escalates; the system adjusts by escalating more, which makes the fraction smaller. |
| Weak oversight | The reviewer is in the loop formally but lacks the evidence, authority, or time to act; the loop is decorative. |
These are the same problems aviation has been working on for forty years, under new acronyms. The orchestrator’s design problem is a version of the cockpit problem, translated into a domain where the controls are sentences and the failure surface is a database row.
4.4.3 Named failure modes for the envelope
The envelope is where those abstract failure modes turn into concrete obligations, because a team that has not named its failure modes has not yet designed its controls. A useful envelope states each mode alongside the artifact that detects it and the action it triggers.
Cosmetic drift. The system begins producing summaries whose tone or formatting has shifted away from the house style. Detection: spot-check a sample of outputs against the style guide each week. Action: tighten the prompt examples; do not change authority.
Interpretive drift. The system begins describing findings in stronger language than the source supports. Detection: compare every output that changes a clinical or regulatory claim against the source record. Action: route the output for human review until the prompt and examples are corrected.
Factual drift. The system begins citing a regulation by the wrong section number, or naming a recipient by an address it inferred from context. Detection: maintain a hard list of fields the system is not allowed to invent; check outputs against that list. Action: strip the field, return a refusal, escalate if it recurs.
Authority drift. The system begins performing an action it was not authorized to perform — closing a ticket, sending a payment confirmation, marking a case resolved. Detection: log every action the workflow takes; review a sample against the contract. Action: revoke authority, rebuild the affected records, run a post-incident review.
Correlated agreement. Three agents use the same model, the same context, and the same flawed source; three agreeing answers are not independent confirmation, they are three versions of the same mistake. Detection: maintain a record of which evaluators share inputs. Action: require an independent second source for any high-consequence decision.
Software engineering met this exact problem in 1978 and got it wrong for a decade. N-version programming, proposed by Chen and Avizienis, was the software analog of redundant hardware: independent teams write separate versions of the same program, the versions run in parallel, and a voter takes the majority — sound, provided the versions fail independently. In 1986 Knight and Leveson tested that provision with twenty-seven independently developed versions and found it false: the failures correlated, because independently written programs made the same classes of mistake on the same hard inputs.29 The field has since relearned the caveat in model form — sample several reasoning chains or several agents and take a majority vote, the technique the research calls self-consistency, and the votes are only as independent as the model and sources behind them.30
Confidence inflation. The system becomes more confident without becoming more accurate — typically because the prompt was tuned for tone or because the evaluation set became easier. Detection: track confidence against calibration on a held-out set. Action: pause the workflow if calibration slips past a threshold.
That list is not exhaustive but it is the minimum a serious envelope should carry, and every team is expected to add the modes specific to its own domain.
4.4.4 Confidence is not consequence
A system can be uncertain and sound certain, and it can be perfectly certain about a fact that has no bearing on the decision — which is why confidence has to be kept separate from consequence. A low-confidence answer in a brainstorming session may be worth exploring; a high-confidence answer that changes a person’s eligibility or a shipment’s destination may still require review.
So record two things separately: how strongly the system supports its claim, and what happens if the claim turns out to be wrong. Those two dimensions route work far better than a single score. High consequence with low confidence demands intervention, and high consequence with high confidence may demand even more scrutiny, particularly when the confidence came from a correlated evaluator or a narrow test set.
4.4.5 Temperature: the dial that helps and does not solve
If variation is the operating condition, the most obvious first move is the temperature setting, and every inference API has one. Temperature reshapes the probability distribution the model samples from: at higher values the model spreads probability across more candidate words, and at lower values it concentrates on the likeliest ones. OpenAI’s own API documentation describes low values as making output “more focused and deterministic.”31 Turn it down and drafts get more consistent; turn it to zero and, in the textbook account, the model always takes the single most likely next word.
Temperature is a real control and a weak one. It narrows sampling variation — the difference between runs caused by the dice roll inside the model — and it costs you something for what it buys, because the variation you are suppressing is the same property that lets a model produce an unexpected option instead of the median one. And the residual variation is not always noise to be minimized: reasoning models deliberately vary their chains of thought between runs, and asking three agents the same question and comparing answers is a working detection technique. The variation is the instrument, not only the hazard.
So treat temperature as a workload decision, not a safety control. A summarizer or a router benefits from running cool — more consistent phrasing, fewer creative departures from the house style, variation pressed toward the cosmetic band. A drafting or research agent benefits from running warmer, because its value is the spread of options it proposes; fence the outputs, do not chill the inputs. What temperature cannot do is make a wrong model right. It reduces the frequency of divergent answers; it does nothing about the wrong answer every run agrees on, and the chapter has already covered that one — correlated agreement is a fence problem, not a dial problem.
Ziad Obermeyer, Brian Powers, Christine Vogeli, and Sendhil Mullainathan, “Dissecting racial bias in an algorithm used to manage the health of populations,” Science 366, no. 6464 (October 25, 2019): 447–453, https://www.science.org/doi/10.1126/science.aax2342. The algorithm, applied to populations of roughly 200 million people per year, used predicted healthcare cost as a proxy for health need; at the same algorithmic risk score, Black patients carried more chronic conditions, and correcting the label eliminated the bias. A preprint is archived at https://www.ftc.gov/system/files/documents/public_events/1548288/privacycon-2020-ziad_obermeyer.pdf.↩︎
Vahid Garousi, “Human Oversight and Overload: Two Hidden and Costly Burdens of AI-Assisted Software Engineering,” arXiv:2606.05770, June 2026, https://arxiv.org/abs/2606.05770. The paper is a review, not original empirical work; it is cited for the cognitive-overload construct and the mechanisms (decision density, context switching, sustained evaluation of plausible alternatives), not for any single quantitative claim.↩︎
Singh, Molloy, and Parasuraman, International Journal of Aviation Psychology (1993), defined automation-induced complacency; Parasuraman and Manzey, Human Factors (2010), review its persistence, https://pmc.ncbi.nlm.nih.gov/articles/PMC6389673. The Air Inter 148 account follows the BEA report summarized by the FAA, https://www.faa.gov/lessons_learned/transport_airplane/accidents/F-GGED: the crew likely selected a 3,300 ft/min descent while intending 3.3 degrees. The chapter does not attribute the accident to complacency alone.↩︎
Convergent taxonomy across multiple 2024–2026 human-in-the-loop reviews: MDPI Entropy 2026 systematic review, https://www.mdpi.com/1099-4300/28/4/377; WJAETS 2024 review, https://wjaets.com/sites/default/files/fulltext_pdf/WJAETS-2024-0012.pdf; 2025 HITL safety-critical survey (Academia.edu), https://www.academia.edu/165688965/Human_in_the_Loop_AI_Systems_A_Review_of_Collaborative_Intelligence_in_Safety_Critical_Applications. The four-mode taxonomy is consistent across reviews; the chapter borrows the taxonomy, not any single study’s specific finding.↩︎
Liming Chen and Algirdas Avizienis proposed N-version programming in 1978 (“N-Version Programming: A Fault-Tolerance Approach to Reliability of Software Operation,” FTCS-8, Toulouse, June 1978, pp. 3–9). J. C. Knight and N. G. Leveson, “An Experimental Evaluation of the Assumption of Independence in Multiversion Programming,” IEEE Transactions on Software Engineering, January 1986, pp. 96–109, tested the independence assumption with independently developed versions and found that failures correlated. A summary of the experiment is at https://www.csc.kth.se/utbildning/kth/kurser/DA2210/vettig13/Seminarier/KnightLeveson.pdf.↩︎
Xuezhi Wang, Jason Wei, Dale Schuurmans, et al., “Self-Consistency Improves Chain of Thought Reasoning in Language Models,” ICLR 2023, https://arxiv.org/abs/2203.11171: sampling multiple reasoning chains and majority-voting the final answers outperforms a single greedy decode. The reliability of the vote depends on the independence of the chains — the N-version caveat, restated for models.↩︎
OpenAI API reference, temperature parameter: “What sampling temperature to use, between 0 and 2. Higher values like 0.8 will make the output more random, while lower values like 0.2 will make the output more focused and deterministic.” https://platform.openai.com/docs/api-reference/chat/create. The same reference recommends adjusting temperature or top_p, not both.↩︎