Treadstone Associates
Definition

What is training data?

Training data is the body of examples an AI model learns from before it is ever used — the text, images, transcripts or records whose patterns the model absorbs during training, which then shape every answer it produces afterward.

Treadstone Associates · Updated 2026

How it’s used in Canada

Canada’s federal, provincial and territorial privacy commissioners, in their joint principles for generative AI, treat what goes into training data as a fairness question, not only a technical one. Developers are expected to work toward “evaluating the training data sets to ensure that they do not replicate, entrench, or amplify historical or present biases – or introduce new biases” — because data drawn from the past can carry the past’s discrimination forward into a system built today.

The same principles make a point that catches most people off guard: training a model on personal information is not the only moment privacy law is engaged. The commissioners state that “the inference of information about an identifiable individual (such as outputs about a person from a generative AI system) will be considered a collection of personal information, and as such would require legal authority”. What a trained model infers about a real person, even something it was never directly told, counts as collecting personal information about them all over again.

Worked example

A firm training a model to flag renewal risk on past mortgage files is not only responsible for how it collected those files in the first place. If the trained model starts inferring something about a specific, identifiable client that was never explicitly recorded — for instance, guessing at a health condition from a pattern of missed payments — that inference is treated as a fresh collection of personal information about that client, needing its own legal basis, independent of how the original training data was gathered.

Related terms

See also: what is a dataset, what is de-identification, what is machine learning.

Where this leads

Deciding what goes into training data, and what is fair to infer from it afterward, is a build decision — custom-ai-solutions covers how that choice gets made and documented before a model ever goes live.