4.2 The failure envelope

Book 2 · The Delegation ContractChapter 4 · section 2 of 5

An envelope, in this book, is the tool an orchestrator uses to define the boundaries of a system that is making a decision — not the paper object that arrives in a mailbox. When a workflow can produce more than one plausible answer, you manage that nondeterminism by defining the set of boundaries and conditions that must be met for a decision to count as valid. Defining that envelope is part of what it means to orchestrate the system.

A physical envelope is binary: a letter is inside or it is not. Envelopes designed for agentic decision-making can define grey areas — regions where the system may draft, recommend, or gather evidence, but may not close the case; where confidence is high enough to proceed with review and too low to act alone; where an unfamiliar documentation pattern triggers neither approval nor denial, but a named human and a training note. Those grey areas are often where the organization learns, where additional training data is collected, where prompts and evaluation sets get revised, and where human intervention is a designed part of the system rather than a failure of it.

Envelopes also come in kinds, and a consequential workflow usually needs more than one. The previous chapter introduced a permission envelope — what an agent may read, change, approve, communicate, spend, or trigger — and an operational envelope — the conditions under which an automated response is expected to be safe. A decision envelope names what the system may decide without a human and which customers, shipments, accounts, or patients that decision can affect. This chapter’s concern is the failure envelope: what variation the system can tolerate, what it cannot, what signals drift, and what evidence causes intervention.

With the right set of envelopes around an agentic topology, nondeterminism can be managed, but it cannot be removed. That follows from composition: if a system contains a component that introduces nondeterministic variation — a sampling model, a concurrent inference batch, a floating-point reduction whose order depends on load — it is nearly impossible to guarantee that the variation will not affect the output.

The research keeps restating that point in different dialects. Setting temperature to zero makes the sampling rule deterministic; it does not make the inference system deterministic, because the result of a request can still depend on batch size, concurrent load, and the order of floating-point reductions.8 Under greedy decoding, changing GPU count, GPU type, or evaluation batch size has been shown to swing a reasoning model’s accuracy by as much as nine percent and its response length by thousands of tokens — not because the prompt changed, but because tiny numerical differences early in the sequence cascaded into different chains of thought.9 Hosted models configured for maximum determinism still fail to deliver identical answers across repeated runs.10 The envelope is what you build because that residue cannot be engineered away under ordinary operating conditions.

The newest wrinkle is that some frontier models will not let you turn it down at all. OpenAI’s reasoning models — the o-series, the GPT-5 waves, and the current GPT-5.6 and GPT-6 lines — fixed their sampling parameters: send a temperature and the API rejects the request, because the vendor wants the chain-of-thought scaffolding to behave consistently and will not let a caller perturb it. The knobs those models expose instead are about effort, not randomness — how long the model thinks before it answers.11 The direction of that change is worth noticing. The industry spent a decade exposing randomness as a dial; its newest reasoning systems took the dial away and made variation a property of the reasoning process itself, which an orchestrator can budget for but cannot turn off.

The fences become a failure envelope when they are written down as one short, named artifact. That document is more useful than a promise that the model will be accurate, because it names the conditions under which the organization has decided the model is allowed to be wrong — and the grey band in which the model is allowed to be unsure without pretending the uncertainty is settled.

A useful envelope answers six questions, in this order:

Question What it forces the organization to decide
What may vary? The bands of variation that are considered within scope (phrasing, ordering, examples, illustrative emphasis)
What must never vary? The invariants the system is not authorized to break (a recipient, a rule, a status, an amount, a protected characteristic)
What signals drift? The measures that, if they move, mean the envelope may no longer fit (a category that should not exist, a metric that should not change)
What requires a second opinion? The thresholds at which the workflow pauses for human review (confidence below a floor, evidence that conflicts, an unusual input class)
What action is reversible? The set of actions whose consequences can be undone in time, and the path that undoes them
What action is not? The set of actions whose consequences cannot be undone, and the gate that prevents them

4.2.1 How envelopes are starting to land in software

Until recently, most of what this chapter calls an envelope lived in runbooks, architecture reviews, and the private judgment of whoever happened to be on call. That is changing, though only just, and the practice remains young. Vendors and open-source projects are inventing places to put the rules — inside prompts, in validators that re-ask or reject bad structure, in cloud policy layers, and in deterministic engines sitting entirely outside the model — while the vocabulary stays unstable, with rails, guards, tripwires, validators, intervention points, and control specifications all naming overlapping ideas. None of it yet amounts to a settled standard for operational orchestration. What it does show is that the industry has started treating “what the system may do when it is unsure” as something that can be declared and checked rather than hoped for in a system prompt.

Five approaches are visible now, each encoding a different slice of the same problem.

Programmable rails around the model. NVIDIA’s NeMo Guardrails treats guardrails as programmable controls between application code and the model — input, dialog, output, retrieval, and execution rails defined in configuration, including Colang dialog flows that can force a path, refuse a class of request, or require a structured extraction before the next step runs.12 That is close to writing an acceptable-outcome range and part of an intervention path into the runtime itself.

Validators on structured output. Guardrails AI popularized a sharper cut at the text boundary: specify the expected structure and quality criteria for a model output, attach validators, and declare what happens on failure — reask, filter, or fix.13 A RAIL-style specification is a miniature envelope for representation: what may appear, what must type-check, and what the workflow does when the model wanders. It does not by itself decide whether a claim may be denied; it decides whether the answer is fit to enter the next stage.

There is also a layer below all of these that works by preventing part of the variation rather than checking for it. Constrained decoding — the mechanism behind OpenAI’s Structured Outputs, introduced in August 2024 — restricts the model at sampling time so the output is forced to match a schema the developer supplies: “model-generated outputs will exactly match JSON Schemas provided by developers,” in the company’s announcement, which reported 100 percent schema adherence for the announcing model against below 40 percent for the previous generation.14 A schema fence buys certainty about the form of the answer — the claim form always comes back well-formed. The values inside the form still come from the model. The structure can be guaranteed. The contents cannot.

Cloud guardrail policies. Managed platforms have started shipping the same idea as service configuration. Amazon Bedrock Guardrails lets an application define denied topics, sensitive-information filters, and contextual grounding checks that score whether a response is grounded in a supplied source and relevant to the query — a practical encoding of “what must never vary,” and a partial answer to factual drift.15 Azure AI Content Safety similarly separates ordinary harm categories from Prompt Shields aimed at jailbreaks and indirect prompt injection in documents the system retrieves.16 These are useful fences. They are also usually shaped around content safety; they do not automatically know a regional routing bias or a claims-authority contract.

Runtime checks on agent loops. Agent SDKs are putting checks where an orchestrator actually intervenes: before the expensive model runs, after the final answer, and around each tool call. OpenAI’s Agents SDK documents input, output, and tool guardrails, tripwires that halt a run, and human-approval interruptions that pause before a side effect until someone accepts or rejects it; the same platform’s evaluation guidance pushes teams from inspecting one-off traces toward graded end-to-end workflow evaluation.17 The intervention path and the evaluation fence are becoming product features rather than custom glue.

Deterministic policy outside the model. The policy engines surveyed in the toolkit chapter already make it possible to move enforcement outside the model; the newer work packages that move for the agent loop itself. Microsoft’s Agent Control Specification, part of its Agent Governance Toolkit, describes a fail-closed, deterministic policy decision runtime: the host builds a JSON snapshot at intervention points across the agent loop, a policy engine returns allow, deny, warn, escalate, or transform, and the host enforces the verdict.18 The architecture argument is explicit — once agents call tools, governance scattered across prompts and framework hooks is not enough; permission and failure rules need a portable artifact that can be evaluated the same way every time. The move has now reached silicon. NVIDIA’s Open Agent Safety Platform, announced September 28, 2026, pairs OpenShell — an open-source sandboxed runtime whose operator policy is verified before the agent runs — with Sentry, a watchdog on BlueField-4 DPUs that monitors agent activity out of band and quarantines an agent that crosses its boundary in milliseconds.[^ch08_oasp] Enforcement in hardware is the strongest form of the placement argument — the check lives where the model’s reach does not go — and it inherits the same limitation as every fence in this section: it enforces the boundary it is given, and cannot tell you whether the boundary is the right one.

One more layer belongs here because of where it sits: outside the model, checking the meaning. Amazon’s Automated Reasoning checks, announced at the end of 2024 as part of Bedrock Guardrails, encode an organization’s rules as formal logic and validate model responses against them with a solver. The documentation is blunt about what separates this from asking a second model to check the first: “Automated Reasoning tools are not guessing or predicting accuracy. Instead, they rely on mathematical proofs to verify.”19 The use cases AWS itself names are regulated ones — insurance eligibility, mortgage approvals, employee benefits — consequential workflows where being plausibly wrong is not an acceptable outcome. This is a factual fence that does not depend on a model’s opinion of another model’s output, which is why it is worth knowing about even for teams that never touch AWS.

The harder observation is that most of these first-generation tools still do not rise to the challenge this chapter is defining. They validate that a run produced an acceptable output — the right schema, a grounded sentence, a denied topic blocked, a tool call paused for approval. That is useful. It is also a narrower problem. What this chapter asks for is validation that the system has the right constraints: the right envelope in which to operate, the right drift signals, the right second-opinion thresholds, the right list of irreversible actions, and the right evaluation set for the people the average metric leaves out. Checking whether a response looks good inside an assumed boundary is not the same as checking whether the boundary itself is adequate for the work.

We have not yet seen tooling that treats the failure envelope — the six questions, the authority map, the grey band, the operating record — as a first-class artifact that can be authored, tested, versioned, and enforced the way a schema or a CI gate is enforced today. Until that arrives, the orchestrator still has to answer the six questions by hand, then map each answer onto whichever layer can currently enforce a slice of it — a prompt rail for dialog bounds, a validator for structured fields, a grounding check for source fidelity, a tool guardrail or policy engine for irreversible actions, and an evaluation set for the variation the dashboard will otherwise miss.


  1. Horace He / Thinking Machines Lab, “Defeating Nondeterminism in LLM Inference,” September 10, 2025, https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/. The load-bearing claim for this chapter is compositional: even when kernels and a forward pass can be described as deterministic in isolation, lack of batch invariance means concurrent load (and thus batch size) can change an individual request’s result. Temperature zero makes the sampling rule deterministic; it does not make the serving system deterministic from the user’s perspective. The post also shows that batch-invariant kernels can restore run-to-run reproducibility under controlled conditions — at a performance cost, and still as an engineering choice rather than a default property of production agentic systems.↩︎

  2. Jiayi Yuan, Hao Li, Xinheng Ding, et al., “Understanding and Mitigating Numerical Sources of Nondeterminism in LLM Inference,” NeurIPS 2025, https://arxiv.org/abs/2506.09501. Cited for the finding that, under bfloat16 with greedy decoding, changing GPU count, GPU type, or evaluation batch size produced up to roughly 9% accuracy variation and multi-thousand-token length differences in a reasoning model (DeepSeek-R1-Distill-Qwen-7B), traced to floating-point non-associativity. The paper proposes mitigations (e.g. LayerCast); the chapter cites it as evidence that numerical residue propagates into outputs, not as proof that no mitigation exists.↩︎

  3. Berk Atıl, Sarp Aykent, Alexa Chittams, et al., “Non-Determinism of ‘Deterministic’ LLM System Settings in Hosted Environments,” Eval4NLP 2025, https://aclanthology.org/2025.eval4nlp-1.12.pdf. Reported accuracy variations up to about 15% across repeated runs of hosted models configured for maximum determinism; none of the evaluated LLMs consistently delivered identical outputs across ten runs. Cited for the operational claim that “deterministic settings” are not a guarantee of identical decisions in hosted agentic systems.↩︎

  4. OpenAI reasoning models (the o-series and the first GPT-5 releases) fix their sampling parameters and reject non-default values — the community-documented API error is “‘temperature’ does not support 0.2 with this model. Only the default (1) value is supported” — and Microsoft’s Azure OpenAI reasoning-model documentation confirms the reasoning lineup does not accept the sampling parameters the chat-completions world grew up with, exposing reasoning_effort (none, minimal, low, medium, high, xhigh, max, varying by model) instead. https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/reasoning↩︎

  5. NVIDIA NeMo Guardrails, open-source toolkit documentation and repository, https://github.com/NVIDIA-NeMo/Guardrails. Cited for the programmable-rails model: input, dialog, output, retrieval, and execution rails between application code and the LLM, including Colang dialog configuration. The citation is for the shape of the control layer, not for any claim that NeMo alone constitutes an operational failure envelope.↩︎

  6. Guardrails AI, RAIL / validator documentation, https://github.com/guardrails-ai/guardrails/blob/main/docs/how_to_guides/rail.md. Cited for the pattern of specifying output structure, field-level quality criteria, and on-fail corrective actions (reask, filter, fix). Used as an example of validation-at-the-text-boundary, not as an endorsement of a particular product or as a complete authority model for tool-using agents.↩︎

  7. OpenAI, “Introducing Structured Outputs in the API,” August 6, 2024, https://openai.com/index/introducing-structured-outputs-in-the-api. The mechanism is constrained decoding, and OpenAI credits its open-source lineage (outlines, guidance, instructor, jsonformer, lark). What the schema guarantees is structure, not the truth of the values inside it.↩︎

  8. Amazon Web Services, “Amazon Bedrock Guardrails,” https://docs.aws.amazon.com/bedrock/latest/userguide/guardrails.html, and “Use contextual grounding check to filter hallucinations in responses,” https://docs.aws.amazon.com/bedrock/latest/userguide/guardrails-contextual-grounding-check.html. Cited for denied topics, sensitive-information filters, and contextual grounding/relevance checks against a supplied source. These are managed content and grounding fences; the chapter does not treat them as a substitute for domain-specific authority contracts.↩︎

  9. Microsoft Learn, “Prompt Shields in Microsoft Foundry,” https://learn.microsoft.com/en-us/azure/ai-services/openai/concepts/content-filter-prompt-shields, and Azure AI Content Safety jailbreak / Prompt Shields documentation, https://learn.microsoft.com/en-us/azure/ai-services/content-safety/concepts/jailbreak-detection. Cited for the separation of user-prompt attacks from document/indirect injection attacks in retrieved or third-party content. Cited as an example of a cloud safety fence at the prompt boundary, not as a complete orchestration envelope.↩︎

  10. OpenAI Agents SDK, “Guardrails,” https://openai.github.io/openai-agents-python/guardrails/; OpenAI API docs, “Guardrails and human review,” https://developers.openai.com/api/docs/guides/agents/guardrails-approvals; and “Evaluate agent workflows,” https://developers.openai.com/api/docs/guides/agent-evals. Cited for input/output/tool guardrails, tripwires, human-approval interruptions before tool side effects, and trace grading / dataset evals for end-to-end agent workflows. Product surfaces change; the load-bearing claim is the placement of checks around the agent loop, not the permanence of any one SDK API.↩︎

  11. Microsoft Agent Governance Toolkit, Agent Control Specification (ACS) policy-engine README, https://github.com/microsoft/agent-governance-toolkit/tree/main/policy-engine. Cited for the fail-closed, deterministic policy-decision pattern: host-built snapshots at intervention points across the agent loop, and normalized verdicts (allow, deny, warn, escalate, transform). Spec versions are explicitly early (e.g. 0.3.1-beta in the published materials); the chapter cites the architectural move — policy outside the model — not a claim that ACS is a finished industry standard.↩︎

  12. Amazon Web Services, “What are Automated Reasoning checks in Amazon Bedrock Guardrails?” https://docs.aws.amazon.com/bedrock/latest/userguide/guardrails-automated-reasoning-checks.html, and “Prevent factual errors from LLM hallucinations with mathematically sound Automated Reasoning checks,” AWS News Blog, December 2024. The checks encode an organization’s rules as formal logic and validate responses with SMT solvers; the quotation is from the blog post. AWS’s named use cases are regulated ones: healthcare, human resources, financial services, insurance eligibility.↩︎