Treadstone Associates
Article · 8 min read

What happens when you send a prompt

It looks instantaneous: type a question, press enter, watch words appear. Underneath, a request goes through several distinct steps before a single word comes back — and the step in the middle, where the actual answer gets produced, works nothing like a lookup or a search. Canada's Cyber Centre describes the starting point plainly: “To create content, LLMs are provided a set of parameters (for example, a query or prompt).” Understanding what happens between that prompt going in and an answer coming back is the fastest way to make sense of why these tools behave the way they do. (Canadian Centre for Cyber Security, ITSAP.00.041)

Treadstone Associates · Updated 2026

Key takeaways

  • • Your words are broken into “tokens” — pieces of words, not always whole words — before the system does anything else with them.
  • • Everything the system can see is assembled into one package first: your prompt, any earlier messages, and any instructions the product added invisibly. If it isn't in that package, the model cannot see it.
  • • The model does not look up an answer. It generates a reply one token at a time, each one chosen from a probability distribution over what could plausibly come next.
  • • Nothing is retrieved from a stored answer bank unless the product deliberately built a retrieval step into the wrapper around the model.

Step one: your words become tokens

A model does not read English the way a person does. The text you type is first broken into tokens — small chunks that are sometimes a whole word, sometimes a fragment of one, and sometimes a single punctuation mark. This is a mechanical, deterministic step: the same text always breaks into the same tokens, and it happens before the model does any “thinking” at all. It matters because the model's actual unit of work is the token, not the word or the sentence — a distinction that becomes important later in this series, particularly when it comes to why these systems struggle with arithmetic.

Step two: everything visible gets assembled into one package

Before generating anything, the system assembles a single package called the context window: your current message, relevant parts of the earlier conversation, and — almost always invisibly to you — a set of instructions the product added on top, telling the model how to behave, what tone to use, or what it is and is not allowed to do. Anthropic's own developer documentation describes this instruction layer as steerable through a “system prompt” that shapes how the model responds to everything that follows it. (Anthropic, tool use documentation) This is the step that explains why the same question, asked inside two different products built on the same underlying model, can produce noticeably different answers: the visible question is identical, but the invisible package built around it is not.

Step three: the model predicts, one token at a time

This is the step most people picture incorrectly. The model is not searching a database of stored answers and returning the closest match. It generates its reply one token at a time, and at each step it is doing the same fundamental operation: given everything so far, what token is most likely to come next? Google DeepMind's own technical description of this process, written to explain how its SynthID watermarking works, puts it plainly: “Large language models generate text one word (token) at a time. Each word is assigned a probability score, based on how likely it is to be generated next.” Its own example: for a sentence like “My favourite tropical fruits are mango and…”, the word “bananas” would score higher than the word “airplanes.” (Google DeepMind, SynthID) That single mechanism — score every possible next token, then pick one — repeated tens or hundreds of times, is the entire generation process. There is no separate step where the model checks its answer against a fact database unless a product has deliberately built one in.

Step four: the chosen tokens stream back as text

Each generated token is converted back into readable text and sent to your screen, which is why longer responses often appear to type themselves out rather than arrive all at once — the system is streaming each token as it is produced, not holding the full answer until it is finished. Generation stops when the model produces a designated end token, hits a length limit set by the product, or is cut off by a rule in the wrapper around it.

A worked example: one prompt, two products, two different answers

Suppose an employee types the identical sentence — “Summarize this contract clause and flag anything unusual” — into two different products that happen to be built on the same underlying model. Product A's wrapper silently adds a system prompt instructing the model to answer as a cautious legal-adjacent assistant that always recommends professional review. Product B's wrapper adds no such instruction and simply forwards the sentence. Both requests reach the same model, tokenized the same way. The context window each one assembles is different, because each product added a different invisible layer on top of an identical visible question — and that is enough, on its own, to produce two visibly different answers with no change to the model at all. This is why comparing “which AI tool is better” by testing the same prompt in two different apps is comparing two different context windows, not necessarily two different models.

Why this is not the same as running a traditional program

A traditional piece of software follows a fixed branch of code for a given input: the same input reliably takes the same path and produces the same output, because a person wrote each decision point explicitly. Nothing in the prompt-to-answer sequence above works that way past the tokenization step. The forward pass through a model is a statistical estimation, and — as the next article in this series covers in detail — the token-selection step usually involves a genuine element of chance, which is part of why asking the identical question twice does not reliably produce the identical answer. Treat the sequence in this article as a request's mechanical path, not as a guarantee about what comes out the other end.

Where retrieval fits in, when it is there at all

Some products add an extra step before any of the above: searching a document store, a website, or a company's own files for relevant material, then inserting what it finds into the context window before the model generates anything. This is a deliberate engineering addition — often called retrieval-augmented generation — not something every AI product does by default. A plain chat interface with no such feature switched on has nothing to retrieve from; it only has what was in its training data and whatever you put directly into the conversation.

Related: what people mean by an AI “model”, why AI answers the same question differently, and, on connecting a model to your own documents and systems, the custom AI solutions hub.

Common questions

Does the model “read” my whole conversation history every time?

Only what fits inside the context window the product assembles for that request, and only what the product chooses to include. A very long conversation can exceed that window, in which case older messages may be summarized or dropped entirely — which is a common, and usually invisible, cause of a model appearing to “forget” something you said earlier.

If it predicts one token at a time, how does it plan out a whole structured answer, like a numbered list?

It does not plan in the way a person outlines an essay before writing it. Structure emerges because a token like a list number or a heading is, in context, a highly probable next token once the surrounding pattern makes that structure likely — the same token-by-token mechanism produces both plain prose and structured output.

Is a search engine doing the same thing as this?

No. A search engine retrieves and ranks documents that already exist. A model without a retrieval feature switched on is not looking anything up at the moment you ask — it is generating new text based on patterns learned during training, which is a fundamentally different operation even when the visible result looks similar.

Keep going in the Academy

This is one page in a plain-English series on how AI actually works and where it fits in a Canadian business.

Back to the AcademyCustom AI solutions