Treadstone Associates
Article · 7 min read

How large language models work

Large language models are trained once and used many times, and the two phases are not the same thing. What gets built during training, and what changes — and doesn't — every time you use it.

Treadstone Associates · Updated 2026

Key takeaways

  • • An LLM is a large set of trained numerical parameters tuned to predict the next piece of text — not a database of stored facts.
  • • Training and the context window are different: Anthropic's own documentation calls the context window a “working memory”, explicitly distinct from “the large corpus of data the language model was trained on”.
  • • Fine-tuning, retrieval and prompting are three different levers — only fine-tuning is further training that changes the model's weights.
  • • NIST's generative-AI risk profile ties confidently wrong output directly to this training mechanism, not to a separate malfunction.

“Large” refers to the training, not to a library it can search

A large language model is not a very big encyclopedia with a fast search bar. It is a very large set of internal numerical settings — usually called parameters, or weights — that were adjusted, over and over, on an enormous quantity of text, so that the model got progressively better at one specific, narrow task: given the text so far, predict what comes next. Nothing in that training process asks the model to store facts as facts. It asks the model to get better at prediction, and a great deal of real-world knowledge gets absorbed as a side effect of text about the world being predictable in patterned ways.

The U.S. National Institute of Standards and Technology (NIST) describes the resulting mechanism directly, in its generative-AI risk profile: (NIST AI 600-1, Generative AI Profile — United States) these systems “generate outputs that approximate the statistical distribution of their training data”. Canada’s federal, provincial and territorial privacy commissioners point to the same NIST material — via a footnote attached to their own accuracy principle — when they require an organization to (OPC, Principles for generative AI) “evaluate the validity and reliability of the generative AI tool for the intended purpose”. Approximating a distribution and being reliably correct are not the same claim, which is the entire reason that principle exists.

Training happens once (mostly); the context window happens every time

It is easy to conflate two different things a model has access to: what it learned during training, and what you put in front of it right now. Anthropic’s own developer documentation for its Claude models draws the line precisely: (Anthropic, Claude API documentation — context windows) the “context window” “refers to all the text a language model can reference when generating a response, including the response itself. This is different from the large corpus of data the language model was trained on, and instead represents a ‘working memory’ for the model.” Training is what shaped the weights before you ever opened the chat window. The context window is the specific, temporary slice of text — your messages, any documents you supplied, the model’s own earlier output — the model is reasoning over for this particular reply. A separate explanation of what that window limits, and what happens at its edge, is worth reading on its own — see what a context window actually limits.

This distinction matters for a practical reason: telling a model something in your prompt does not teach it anything permanently. It only affects this conversation, because the weights themselves — the actual “training” — are not being changed by anything you type.

Three different ways to change what a model does, and only one of them is “training” it further

Anthropic’s developer glossary defines the three levers a business is actually choosing between when it wants a model to behave differently, and the definitions are worth reading exactly as written because the three are frequently used loosely as if they were interchangeable: (Anthropic, developer glossary) fine-tuning is “the process of further training a pretrained language model using additional data. This causes the model to start representing and mimicking the patterns and characteristics of the fine-tuning dataset.” That is real, additional training — it changes the weights. Retrieval-augmented generation (RAG), by contrast, “combines information retrieval with language model generation… a language model is augmented with an external knowledge base or a set of documents that is passed into the context window” at the moment of the request — the weights never change; the model is simply handed more to read. A full explanation of that mechanism is its own page — see how retrieval grounds an AI answer. The third lever, prompting, changes neither the weights nor the knowledge base — it only changes the instructions and framing inside the context window for this one exchange.

The practical business question is rarely “which of these is best” in the abstract. It is which lever actually reaches the problem you have: a knowledge-freshness problem is usually a retrieval problem, a tone-and-format problem is usually a prompting problem, and a durable behaviour change across an entire product is one of the few situations where fine-tuning is the right tool — and even then, it is additional training on top of an already-trained model, not a replacement for it.

What this design cannot promise, structurally

Because the underlying mechanism is prediction from a learned statistical pattern rather than lookup from a verified record, a model has no internal signal that distinguishes a fact it has seen stated reliably many times from a continuation that merely sounds like the others it was trained on. NIST’s risk profile names the resulting failure mode (NIST AI 600-1 — United States) “confabulation” — “the production of confidently stated but erroneous or false content” — and traces it to the same training mechanism described above, not to a separate bug. The practical takeaway for a business is not that the technology is unreliable across the board; it is that reliability tracks how well-represented a topic was in training and in whatever is currently in the context window, not how confident the output sounds.

None of this is unique to any one company's models — it is a property of training a model this way at all, regardless of which organization did the training or how large the resulting model is. A larger model trained on more text can absorb more patterns, but scale changes how much the model has learned to predict, not whether prediction is fundamentally what it is doing. That is why the mechanism described here is worth understanding on its own terms, separately from any specific model's marketed capabilities, which change release to release while the underlying mechanism does not.

Related, within this hub: how ChatGPT generates an answer, what a context window actually limits, and what open-weight AI models change. For how this bears on build-versus-buy decisions, see our approach to custom AI solutions.

Common questions

Is a large language model the same thing as a search engine?

No. A search engine retrieves and ranks existing documents. A base language model generates new text by predicting likely continuations from patterns learned during training, with no guaranteed link back to a specific source unless a separate retrieval step is added on top.

Does a bigger model simply know more facts?

More training data and parameters generally mean more patterns get absorbed, but that is not the same as guaranteed correctness on any one question. NIST's own generative-AI risk profile ties confidently wrong output directly to how these models are designed, regardless of scale.

What is the difference between training and fine-tuning?

Training builds the base model from a very large, general body of text. Fine-tuning is additional training layered on top of an already-trained model, using a smaller, more specific dataset, so the model starts mimicking the patterns of that narrower dataset.

Working out what a model like this can actually carry for your business?

A short call is enough to map the mechanism against what you need it to do.