Treadstone Associates
Article · 9 min read

Fine-tuning vs prompting vs retrieval

All three change how a model answers. Only two of them change what it actually knows. Picking the wrong one is a common, expensive way for a business to solve a problem it does not have while leaving the one it does have untouched.

Treadstone Associates · Updated 2026

Key takeaways

  • • Prompting changes instructions, not knowledge — it shapes how the model uses what it already has, and nothing more.
  • • Fine-tuning changes the model's own internal parameters, which is durable but goes stale exactly like the original training data does.
  • • Retrieval hands the model a document at answer time, direct from a source you control, without touching the model itself.
  • • Retrieval is the only one of the three that can honestly hand back a link you can open and check.

Prompting changes instructions, not knowledge

Prompting — the instructions and examples you give a model at the moment you ask it something — shapes how it uses what it already has: tone, format, which of its existing knowledge to emphasize for this particular question. It does not add a single fact the model was not already exposed to during training, and it does not persist: the next conversation starts from the same base model with no memory of the instructions you gave last time, unless you supply them again.

This makes prompting the right tool for a narrow class of problem: the model already has the knowledge, and the issue is presentation, format or emphasis rather than a missing fact. Asking a model to answer in point form instead of paragraphs, or to adopt a specific tone, is a prompting problem. Asking it to know something about your business that it was never trained on is not, no matter how carefully the prompt is worded.

Fine-tuning changes the model itself

Fine-tuning trains the model further on your own examples, adjusting its internal parameters so its default behaviour shifts — consistently answering in your firm's format, for instance, or getting measurably better at one narrow, repeated task. Canada's Cyber Centre lists this as a standard practice for anyone operating a generative system: continuously fine-tune or retrain the AI system with appropriate external feedback to improve the quality of outputs. That framing is accurate about the trade-off as well as the benefit — fine-tuning is comparatively expensive and slow to redo, so a fine-tuned model's knowledge goes stale in exactly the way its original training data did, just on your own schedule instead of the base model provider's.

Fine-tuning is best suited to changing how a model behaves consistently — its style, its default assumptions, its skill at one repeated narrow task — rather than to keeping it current on facts that change often. A model fine-tuned to always format a quote the way your firm formats one will keep doing that reliably. A model fine-tuned last quarter on your product catalogue will not automatically know about a product added last week; that gap is what the next approach exists to close.

Retrieval changes what the model can see at answer time

Retrieval takes a different approach entirely: instead of baking information into the model's parameters, it hands the model a specific document at the moment it answers, pulled live from a source your business controls. This is the direct fix for a limitation Canada's privacy principles flag explicitly: a provider should disclose where the training dataset is time bounded (i.e. only contains information up to a certain date). Retrieval sidesteps that limitation by design — the model is answering from whatever is in the store right now, not from whatever was true when it was last trained.

Retrieval requires its own maintenance, just of a different kind than fine-tuning: someone has to keep the document store itself current, organized and searchable. It trades a slow, expensive retraining cycle for an ongoing, usually cheaper, curation task — and unlike fine-tuning, updating the store takes effect on the model's very next answer, with no retraining step in between.

Cost, staleness and control: the three-way trade-off

Fine-tuning has the highest setup cost and produces the most durable behaviour change, but it must be redone as your material changes. Prompting is the cheapest and most flexible of the three, and the weakest at controlling what the model actually knows. Retrieval needs a searchable, current store of your own documents to work at all, but keeps answers current without retraining — and it is the only one of the three that can honestly hand back a citation a person can open and check, which is the exact gap covered in can AI cite its sources reliably.

A worked example: a firm's own policy manual

Consider a business that wants an assistant to answer staff questions about its internal policy manual accurately. Prompting alone will not work if the manual was never part of the model's training data — no amount of clever instruction gives the model facts it has never seen. Fine-tuning on the manual would teach the model its contents at the moment of training, but the moment the manual is revised, the fine-tuned model keeps confidently repeating the old version until it is retrained. Retrieval, by contrast, means the assistant looks up the current manual text at the moment of each question and answers from what it just read — so a policy change takes effect for every future question as soon as the document itself is updated, with no retraining step at all. For a document that changes even occasionally, retrieval is usually the mechanism that matches the actual problem.

Whichever you choose, it is still a new component to track

The NIST framework's own risk-mapping step requires that approaches for mapping AI technology and legal risks of its components – including the use of third-party data or software – are in place, followed, and documented. A fine-tuned model, a retrieval index and a prompt template are each, in that sense, a new component: whose data trained the fine-tune, whose documents populate the index, whose examples shaped the prompt. Choosing one of the three does not remove the need to know the answer to that question — it just changes where the answer lives.

See also can AI cite its sources reliably and generative AI vs agentic AI.

Common questions

If we fine-tune a model on our documents, does it know today's version of them tomorrow?

No. A fine-tuned model only knows what it was trained on at the point the fine-tuning ran. A document that changes the next day needs either a fresh round of fine-tuning or, more practically, a retrieval layer that reads the current version directly.

Is prompting a cheaper substitute for fine-tuning?

No, they do different jobs. A well-written prompt cannot teach a model a fact it was never exposed to during training — it can only change how the model presents what it already has.

Can a business combine all three?

Yes, and many practical tools do: a fine-tuned model for a consistent in-house voice, retrieval for facts that need to stay current, and prompting to adapt the same underlying setup to each specific task.

Which of the three should a business try first?

Usually prompting, because it costs nothing to test and rules out the cheapest explanation first. If a well-written prompt genuinely cannot get the model to behave the way you need, that is the signal to look at retrieval or fine-tuning next, and specifically at whether the gap is about missing knowledge (retrieval) or inconsistent behaviour (fine-tuning).

Does retrieval replace the need for any fine-tuning at all?

Not always. Retrieval solves the knowledge-freshness problem well, but it does not, on its own, make a model consistently adopt a specific tone or format the way fine-tuning can. The two solve different halves of the same broader goal, which is exactly why many practical setups use both.

Where this goes next

Choosing between these three is usually part of a larger build-vs-buy decision. For that, see scoping a custom AI assistant, from idea to working prototype.