Treadstone Associates
Definition

What is inference in AI?

Inference is what a model does after training is finished: it takes a new input — a document, a question, a prompt — and produces an output, without changing anything about the model itself in the process. No Canadian regulator or court uses the word “inference” this way in the material available here; the distinction it names is nonetheless the one Canada’s own AI guidance draws, just without that label.

Treadstone Associates · Updated 2026

How the term is used in Canada

Canada’s Cyber Centre describes the training half of the split without naming it: “Generative AI is a type of AI that generates new content by modelling features from large datasets that were fed into the model.” That happens once, when the model is built. The other half happens every time afterward, whenever the finished model is put to work, which the same guidance also describes without naming it: “To create content, LLMs are provided a set of parameters (for example, a query or prompt).” Nothing about the datasets that originally shaped the model is being re-read at that second stage — the model is applying what it already learned to the one input in front of it. That second stage is inference.

The same before-and-after line shows up again in Canada’s privacy guidance, for an unrelated reason. Its principles separate a “Developer or Provider”, who determines “how it is initially trained and tested”, from an organization “using a generative AI system as part of their activities”. For most Canadian businesses running a model someone else built, “using” means running inference — sending prompts and reading answers — and nothing about that activity retrains the model.

Worked example

A lender’s underwriting tool was trained months ago on a large set of past decisions. Every time a new application arrives and the tool scores it, that scoring run is inference — the model’s own settings do not change because one more file passed through it. If the lender later decides the tool is drifting and deliberately retrains it on a fresh batch of files, that separate, deliberate step is fine-tuning, not inference. The two are easy to conflate because both involve feeding the model data; the difference is whether anything about the model is meant to change afterward.

See also fine-tuning, foundation model and large language model.

Keep going in the Academy

This is one term in a plain-English glossary on how AI actually works and where it fits in a Canadian business.

More definitions Custom AI Solutions