Type a question into a chat assistant and a reply appears a few words at a time. Nothing is being looked up. The model is running the same short calculation once per token, for as long as the reply lasts.
Key takeaways
A chat assistant does not compose a reply the way a person drafts an email, planning the whole thing and then writing it down. It builds the reply piece by piece, choosing one small chunk of text — a token, which is often a word, part of a word, or a punctuation mark — and then repeating the same narrow calculation to choose the next one. Google DeepMind’s own technical description of how its models write text states the mechanism plainly: (Google DeepMind, SynthID) “Large language models generate text one word (token) at a time.” That sentence, written to explain a watermarking tool, is actually the cleanest public description of how any large language model produces text at all — watermarking simply nudges the probability scores behind each word after the fact.
This matters for how you read an answer. There is no stored sentence sitting in a database waiting to be displayed. Every reply, including one you have seen the model give before, is assembled fresh, token by token, at the moment you ask.
At each step, the model takes everything written so far — your question, any earlier turns in the conversation, and every token it has generated in its own reply up to that point — and runs it through a fixed calculation that produces one thing: a probability score for every token in its vocabulary, representing how likely that token is to come next. The model then selects one. Sometimes it simply takes the highest-scoring option; more often a controlled amount of randomness is allowed into the choice, which is a separate, adjustable setting rather than an accident, and is worth its own explanation further down this hub. Whichever token is chosen is appended to the sequence, and the entire calculation runs again from scratch to choose the next one.
This is why a long answer can start strong and drift: nothing about the process changes, but the sequence the model is reasoning over keeps growing, and a small early choice — a word, a framing, a stray assumption — becomes part of the fixed context every later token is chosen against.
Because each token is chosen for being statistically probable given everything before it, a model can produce a sentence that reads as confident and well-formed while stating something false. The U.S. National Institute of Standards and Technology (NIST) names this outcome in its generative-AI risk profile and traces it directly to the generation mechanism just described: (NIST AI 600-1, Generative AI Profile — United States) “Confabulations are a natural result of the way generative models are designed: they generate outputs that approximate the statistical distribution of their training data; for example, LLMs predict the next token or word in a sentence or phrase.” NIST’s word for the phenomenon is “confabulation”; it is the same thing most people call a hallucination.
Canada’s federal, provincial and territorial privacy commissioners point to the same underlying concern from the governance side rather than the engineering side. Their joint principles for generative AI — which cite the NIST framework directly in support of this point — state that an organization must (OPC, Principles for generative AI) “evaluate the validity and reliability of the generative AI tool for the intended purpose”, adding that “tools must be accurate throughout the intended lifecycle of the tool and across the variety of circumstances in which they are used.” Fluency is not evidence of that; fluency is what the token-by-token mechanism always produces, whether the underlying claim is well-supported or not.
Say you ask a chat assistant to finish the sentence “The capital of Canada is…”. The model does not look up “Ottawa” in a table of facts. It reaches the end of that phrase with a sequence of tokens behind it, computes a probability distribution over every possible next token, and finds that “Ottawa” scores far above alternatives like “Toronto” or “Montreal” — not because it is checking an atlas, but because that continuation appeared overwhelmingly more often than the alternatives in the text it was trained on, in exactly this kind of sentence. It writes “Ottawa,” appends it to the sequence, and the loop continues to decide what comes after that.
Now change the question to something obscure — a niche regulation, a small company’s internal policy, a paper nobody widely cited. The mechanism has not changed at all. But the training data behind that particular continuation is thin or absent, so the probability distribution the model computes is flatter and less reliable, and the token it lands on can be a plausible-sounding invention rather than a fact. The model has no separate step that checks “do I actually know this,” because knowing and not-knowing are not different modes it operates in — they are just different shapes of the same probability distribution.
The token-by-token loop needs a stopping condition or it would never end. Anthropic’s API documentation for its Claude models lays out the actual stop conditions a request can hit, and they map onto two genuinely different situations: (Anthropic, Claude API reference — stop reasons) an end_turn result means “Claude finished its response naturally”, while a max_tokens result means “the response reached your max_tokens limit” before the model was finished — a length cap set by whoever built the product you’re using, not a sign the model ran out of things to say. That second case, and the cut-off replies it produces, has its own explanation elsewhere in this hub.
This is also why a model has no reliable way to tell you, in advance, how long its own answer will be, or to know it is running low on space: it is choosing one token at a time with no plan for the whole reply, so the stop condition is discovered by hitting it, not predicted ahead of it.
None of this makes the mechanism untrustworthy in a blanket sense — next-token prediction over an enormous amount of text produces genuinely useful writing, summarizing and drafting. It does mean the reasonable default is to treat a generated answer the way you would treat a confident colleague speaking from memory on a topic they have not checked recently: useful as a starting draft, and worth verifying before it goes into something that matters, especially where the topic is narrow, numeric, or something the model would have seen rarely in training. Two techniques change what actually goes into that per-token calculation — giving the model real source documents to reason from, which is how retrieval grounds an AI answer, and adjusting how much randomness is allowed into each token choice, covered in what actually changes an AI’s output — and both are worth understanding before you decide how much to trust a given answer.
Content-provenance efforts such as (C2PA) and (Content Credentials) address a related but different problem: not whether a generated answer is accurate, but whether a piece of content can be traced back to having been machine-made at all. They are a labelling layer sitting on top of the generation mechanism described above, not a check on what the mechanism produces.
Related, within this hub: how large language models work and how retrieval grounds an AI answer. If you are weighing whether to build something on top of a model like this rather than just use one off the shelf, see our approach to custom AI solutions.
Not by default. The base mechanism is generation from a probability distribution learned during training, not a lookup against a live source. Some products add a separate retrieval step that fetches real documents before generating — a materially different mechanism, covered in how retrieval grounds an AI answer.
A controlled amount of randomness is normally allowed into which token gets chosen at each step, so two runs of the same prompt can diverge in wording, and occasionally in substance, even when nothing about the setup changed. See what actually changes an AI's output for the specific settings involved.
No. Fluency is what the token-by-token mechanism always produces, because every token is chosen for being statistically probable, not for being verified. NIST's own generative-AI risk profile describes confidently stated but false content as a natural result of how these models are designed, not a malfunction.
A short call is enough to map what the mechanism can and can’t carry for your use case.