The same question can get a different answer twice, and it is rarely a mystery. A short taxonomy of what genuinely changes an AI's output, and what only looks like it did.
Key takeaways
Ask a chat assistant the same question twice and the wording, and occasionally the substance, can come back different both times. The most common cause is not a flaw — it is a deliberately exposed setting. Anthropic’s developer glossary defines it directly: (Anthropic, developer glossary) “Temperature is a parameter that controls the randomness of a model’s predictions during text generation. Higher temperatures lead to more creative and diverse outputs… Lower temperatures result in more conservative and deterministic outputs that stick to the most probable phrasing and answers.” Whoever built the product you are using chose a temperature setting, and that choice — not the question you asked — largely determines how much variation you should expect between runs.
The same glossary entry goes further than most people expect: (Anthropic — developer glossary) “users may encounter non-determinism in APIs. Even with temperature set to 0, the results will not be fully deterministic and identical inputs may produce different outputs across API calls. This applies both to Anthropic’s first-party inference service and to inference through third-party cloud providers.” In plain terms: there is no setting that guarantees an identical answer twice, only settings that make it more or less likely.
Temperature explains run-to-run variation with an unchanged prompt. It does not explain every difference you will see, and three other levers are routinely confused with it. Changing the words of your prompt changes the output because you have changed what the model is reading, not because anything about the model itself moved — the same model, asked differently, reasons over a different context window. Giving the model new documents to read — retrieval — changes the output because the material it is drawing on changed, a mechanism covered on its own at how retrieval grounds an AI answer. And genuinely retraining the model further on new data — fine-tuning — is different again: Anthropic’s glossary defines it as (Anthropic — developer glossary) “the process of further training a pretrained language model using additional data,” which “causes the model to start representing and mimicking the patterns and characteristics of the fine-tuning dataset” — a permanent change to the weights themselves, not a per-request setting.
A fifth cause is easy to overlook because it happens outside anything you control: the provider quietly updates or replaces the underlying model version behind an API endpoint. Nothing about your prompt, your settings or your documents needs to change for the output to shift when that happens — which is one reason a workflow that depends on exact wording is worth re-testing periodically, independent of anything the business itself changed.
It is worth asking why a product would ever want variation instead of always taking the single most probable token. The reason traces back to the same generation mechanism covered in how ChatGPT generates an answer: a model produces outputs that, in the words of the U.S. National Institute of Standards and Technology’s generative-AI risk profile, (NIST AI 600-1 — United States) “approximate the statistical distribution of their training data.” Always taking the single highest-scoring token at every step tends to produce flatter, more repetitive writing — genuinely useful for a task with one clearly correct answer, and noticeably dull for open-ended drafting. Allowing some controlled randomness into the choice is a deliberate trade: less predictability in wording, in exchange for output that reads less like the same template every time. Neither setting is “more correct” in the abstract — they are tuned for different jobs, which is exactly why the setting is exposed as a choice rather than fixed.
A business relying on an AI-generated output for anything repeatable — a customer-facing reply, a compliance summary, a numeric extraction — needs to know which of these levers actually moved before assuming a result is “wrong” or a process is “broken.” A different answer from a higher-temperature setting is expected variation, not a defect. A different answer because the prompt changed is expected, not mysterious. A different answer with an unchanged prompt, unchanged documents, and temperature already at its lowest setting is the one case that is actually worth investigating — and even then, Anthropic’s own documentation is clear that a small amount of unavoidable variation remains even at temperature 0.
Canada’s federal, provincial and territorial privacy commissioners frame the underlying obligation plainly in their joint principles for generative AI, in language written for reliability generally that applies directly to a business depending on consistent output: an organization is expected to (OPC, Principles for generative AI) “evaluate the validity and reliability of the generative AI tool for the intended purpose,” and to treat that evaluation as ongoing rather than one-time, because “tools must be accurate throughout the intended lifecycle of the tool and across the variety of circumstances in which they are used.” A workflow that was tested once, against one model version, on one day, has not actually established that — which is precisely why a silent model update behind an API endpoint, and not only a temperature setting, belongs on the list of things to check when output drifts.
None of this is an argument against using the technology for anything that must be repeatable. It is an argument for building a check into the workflow — comparing outputs against a rule or a second pass, rather than assuming one run is representative — precisely because the mechanism does not promise sameness even when nothing about your side of the request changed.
Related, within this hub: how ChatGPT generates an answer and how large language models work. For workflows where repeatability actually matters, see how we approach custom AI solutions.
Most commonly, a randomness setting called temperature, chosen by whoever built the product. Anthropic's own documentation adds that even at the lowest possible setting, identical inputs may still produce different outputs across separate calls.
It can reduce ambiguity, but clarity and determinism are different things. A precisely worded prompt still runs through the same randomness setting on the generation side, so wording alone does not remove run-to-run variation.
Often the provider has updated the underlying model version behind the same API endpoint. It is worth checking for that before assuming the workflow itself is broken.
A short call is enough to work out which lever actually needs controlling.