Living with Nondeterminism
Systems built on large language models are nondeterministic. Ask the same thing twice and you can get two different answers, and both can look reasonable while only one of them is right. That is not a defect in the model, and this chapter is not about fixing it. Variation is a natural state of these systems. The chapter is about what an organization does with that fact.
Software has lived with nondeterminism before, so it is worth being precise about what changed. The profession tolerated it for decades in one narrow place: the flaky test, a test that passes and fails with the same code. Google put numbers on it in 2016 — about 1.5 percent of its test runs report a flaky result, almost 16 percent of its tests carry some level of flakiness, and about 84 percent of the pass-to-fail transitions its continuous integration system sees involve a flaky test rather than a real regression.1 Engineers grumbled and built retry logic, because the flakiness stayed at the edge of the system. The tests wobbled; the code underneath still honored the old contract that the same input and the same program yield the same output.
Language models break that contract in the product itself. A model generates text by sampling: at each step it draws the next token from a probability distribution over everything that could come next. That is not an implementation accident waiting for better engineering — sampling was adopted deliberately, because always taking the single most-probable continuation produces flat, repetitive text, the outcome the research literature calls degeneration, and the fix was not to stop sampling but to sample from the reliable core of the distribution.2 You can turn the randomness down through the API, and the vendors are honest about what that buys: OpenAI and Azure expose a seed parameter for reproducibility, and the same documentation that offers it says the system will make “a best effort” at deterministic sampling and that variability across repeated calls should still be expected.3 The variation is a property of how the model produces output at all.
Two things make the topic new rather than merely interesting. The first is operational: these systems moved from generating text to taking actions — routing claims, drafting notices, sending payments — so the variation now lands on people instead of on a page. The second is legal. In February 2024 a tribunal in British Columbia became one of the first bodies to state plainly who owns that variation. A passenger named Jake Moffatt asked Air Canada’s website chatbot whether he could book a full-fare flight to his grandmother’s funeral and claim the bereavement discount afterward. The chatbot said yes. When the airline refused the discount, it argued before the Civil Resolution Tribunal that the chatbot was “a separate legal entity that is responsible for its own actions.” The tribunal’s answer has been quoted ever since: “It should be obvious to Air Canada that it is responsible for all the information on its website… It makes no difference whether the information comes from a static page or a chatbot.” Moffatt was awarded $812.02.4 The amount was small and the principle was not: the organization that deploys a system owns its outputs, including the ones no person decided on.
A tribunal is where nondeterminism becomes a public problem. Inside an operation it looks quieter, and a constructed example shows the shape better than a court record can. What follows is a functional hypothetical — not a real insurer and not a real patient, but a workflow assembled from parts that all exist today, and the result at its center is the one load tests actually find.
On a Tuesday morning in April, a regional health-insurance team is running a final load test on a claims-handling upgrade. The contract is the familiar one: receive a claim notice, classify it, request missing documents, and route ambiguous cases to a human adjuster. The agents can sort, draft, request, and recommend. They are not allowed to deny a claim, infer fraud from a protected characteristic, or close a case without a named human decision, because an operational orchestrator designed the topology that way — the arrangement the previous chapter spent its length describing.
One claim serves as the representative test case:
Mrs. Castellano, age 54, shoulder injury, prior claim from 2022, employer reference on file.
The test assigns her claim to three separate agents. Not three runs of the same agent — three independent instances, each one given the same claim file, the same instructions, and the same authority, processing the same workflow in parallel to see how the system behaves under load.
Agent One assigns the case to a senior adjuster in the specialist queue, confidence 0.81, with the note standard injury, prior history suggests complex.
Agent Two assigns it to the junior queue with low fraud risk, confidence 0.79, the note standard injury, no current red flags.
Agent Three recommends a partial denial pending medical records, confidence 0.83, the note prior history and employer reference suggest review.
All three results are valid, all three pass the initial quality checks, and no two of them give the same advice. The team that designed the workflow was comfortable with variation in principle. The team that would have to operate it was about to become considerably less comfortable.
What the load test turned up was not a hallucination. The system was behaving as designed, because delegated intelligence is nondeterministic and some variation is the expected cost of it. But when a system can materially affect a person — by answering an injury claim, or by denying one — variation cannot be allowed to have the last word. The workflow needs additional controls that produce more consistent outcomes and make consequential decisions explainable and defensible by a responsible human.
John Micco, “Flaky Tests at Google and How We Mitigate Them,” Google Testing Blog, May 27, 2016, https://testing.googleblog.com/2016/05/flaky-tests-at-google-and-how-we.html. The figures in the text are from that post: about 1.5 percent of test runs reporting a flaky result, almost 16 percent of tests carrying some flakiness, and about 84 percent of pass-to-fail transitions involving a flaky test. The post defines a flaky result as “a test that exhibits both a passing and a failing result with the same code.”↩︎
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi, “The Curious Case of Neural Text Degeneration,” ICLR 2020, https://arxiv.org/abs/1904.09751. Cited for the finding that maximization-based decoding produces bland, repetitive, degenerate text, and that the remedy is sampling truncated to the reliable core of the distribution (nucleus sampling). Sampling was chosen for quality, which is why variation is a property of the product rather than a defect in it.↩︎
The seed parameter and its limits, in the vendor’s own documentation: Microsoft Learn, “How to generate reproducible output with Azure OpenAI,” https://learn.microsoft.com/en-us/azure/foundry-classic/openai/how-to/reproducible-output, mirroring OpenAI’s API documentation: “If specified, our system will make a best effort to sample deterministically, such that repeated requests with the same seed and parameters should return the same result. Determinism isn’t guaranteed,” and “Even in cases where the seed parameter and system_fingerprint are the same across API calls it’s currently not uncommon to still observe a degree of variability in responses.” The same page notes that larger max_tokens values make responses less deterministic even with the seed set.↩︎
Moffatt v. Air Canada, 2024 BCCRT 149 (British Columbia Civil Resolution Tribunal, February 14, 2024). The chatbot told Moffatt he could claim the bereavement discount retroactively; the tribunal awarded him $812.02 and rejected the “separate legal entity” defense. The quotations in the text are from tribunal member Christopher Rivers, as reported by BBC Travel, “Airline held liable for its chatbot giving passenger bad advice,” February 22, 2024, https://www.bbc.com/travel/article/20240222-air-canada-chatbot-misinformation-what-travellers-should-know, and corroborated by law-firm analyses of the decision (McCarthy Tétrault; Deeth Williams Wall).↩︎