A colleague asks the identical question you asked yesterday, into the identical tool, and gets a noticeably different answer. For software in general this would be alarming — a calculator that gave 9 one day and 9.2 the next would be broken. For a generative AI system, some amount of this is normal, expected, and traceable to a specific, deliberate design choice in how the system picks its words.
Key takeaways
The previous article in this series covered the core generation step: a model scores every plausible next token and picks one, repeated until the answer is complete. Google DeepMind's own technical description of this, written to explain how its SynthID watermarking works, states it directly: “Large language models generate text one word (token) at a time. Each word is assigned a probability score, based on how likely it is to be generated next.” Its own worked example: given “My favourite tropical fruits are mango and…”, the word “bananas” scores higher than the word “airplanes.” (Google DeepMind, SynthID) The critical detail for this article is what happens after the scoring: the system is not required to pick only the single highest-scoring token every time. Depending on how it is configured, it can sample from among several plausible high-scoring options — which is exactly the step that introduces genuine variation between two otherwise identical requests.
Always picking the single most probable next token, every time, produces text that reads as noticeably flat and repetitive — the safest word, every time, in a row. Allowing the system to sample among several plausible options, weighted by their scores, produces more natural, varied language, at the cost of run-to-run consistency. This trade-off is usually exposed as an adjustable setting, often called “temperature,” that a product or a developer can turn up for more creative, varied output or down for more consistent, predictable output. A product built for creative writing and a product built for extracting a structured field from a form are reasonably tuned very differently on this dimension, even if they run on the same underlying model.
The United States' National Institute of Standards and Technology — a foreign standards body, cited here only because the Office of the Privacy Commissioner of Canada's own generative-AI guidance directs readers to it for exactly this topic — frames reliability as the “ability of an item to perform as required, without failure, for a given time interval, under given conditions,” and separately notes that “Accuracy and robustness… can be in tension with one another” in AI systems. (NIST AI RMF, via the OPC's own footnote 16) (OPC, generative AI principles) A system tuned for varied, creative output is, by construction, tuned to be less rigidly repeatable — the two properties are pulling in different directions, and a product's designers are making a deliberate choice about where to sit on that line for a given use case, not failing to achieve consistency by accident.
Not every difference between two answers traces back to sampling randomness, and conflating the causes leads to the wrong fix. If the same question produces a completely different kind of answer, not just a differently-worded version of a similar one, the more likely explanation is one covered earlier in this series: a different system prompt was silently added by the product, the underlying model itself was swapped by the vendor, or the context window contained different information the second time — a longer conversation history, or a document that was or was not attached. Genuine sampling variation produces answers that are different in wording but similar in substance; a different context or a different model can produce answers that disagree on the actual facts, which is a different problem with a different fix.
The right response to run-to-run variation depends entirely on the task, not on the tool in the abstract. For open-ended drafting — brainstorming options, writing a first pass at marketing copy, generating several angles on a problem — variation between runs is close to a feature: it is the mechanism that gives you several genuinely different starting points instead of one. For extracting a specific field from a document, classifying a support ticket into one of five fixed categories, or producing a number that should be stable, the same variation is a liability, and the correct response is not to complain that the model is unreliable in general — it is to check whether the task and the settings are matched to a consistency-first use case, or whether a narrower, more structured way of asking the question would remove the room for the model to wander at all.
Ask a model three separate times to “list three benefits of remote work” with sampling turned on. Run one might lead with flexibility, reduced commuting and cost savings. Run two might lead with improved focus, better work-life balance and a wider talent pool. Both are reasonable, defensible answers to a genuinely open-ended question — there is no single correct list of exactly three benefits, so sampling different plausible options each time is not an error. Contrast that with asking the same model three times for the current federal small business tax rate: here there is one correct figure, and if the three runs disagree on the number itself rather than just the phrasing around it, that is a reliability problem worth investigating, not a benign product of creative sampling.
Related: what happens when you send a prompt, why AI gets simple things wrong, and, on choosing settings and models for a given task, the custom AI solutions hub.
It can usually be reduced substantially by lowering the sampling setting, but most products do not expose a literal “always pick the top-scoring token” switch to an ordinary user, and even a fully deterministic setting does not protect against a factually wrong top-scoring answer — it only makes the same answer repeat consistently, right or wrong.
Not by itself. Variation is a tuning choice suited to some tasks and not others, not a defect. A model tuned for varied, natural-sounding writing is doing what it was built to do; the same setting would be a poor choice for a task that needs an identical structured field extracted the same way every time.
Favour tasks and settings built for consistency over creativity, keep the context window identical between runs, and where a factual figure matters, verify it against a primary source rather than relying on repeated runs of the same prompt to converge on the truth — repetition alone does not make an answer correct.
This is one page in a plain-English series on how AI actually works and where it fits in a Canadian business.