Data minimization is the practice of collecting, retaining and feeding into a system only the personal information genuinely needed for a specific, identified purpose — not the largest dataset that happens to be available.
The principle predates AI and sits directly in PIPEDA’s Schedule 1, Principle 4: “The collection of personal information shall be limited to that which is necessary for the purposes identified by the organization”, with the further instruction: “Organizations shall not collect personal information indiscriminately”. That general rule is echoed in section 5(3)’s appropriate-purposes test: “An organization may collect, use or disclose personal information only for purposes that a reasonable person would consider are appropriate in the circumstances”. Neither provision was written with AI in mind, and both apply to it anyway.
Canada’s privacy commissioners have since restated the principle specifically for generative AI. Their joint principles direct developers and users to: “Consider whether the use of a generative AI system is necessary and proportionate”, adding that this “should be evidence-based and establish that the tool is both necessary and likely to be effective in achieving the specified purpose,” and to prefer anonymized, synthetic or de-identified data over personal information wherever the personal information is not actually required. Put together, the federal rule is one general test — necessary for an identified purpose — applied twice: once to what an organization collects in the first place, and again, by the commissioners’ AI-specific principles, to what it feeds into a generative AI system afterward.
A brokerage wants an AI tool to draft renewal reminder emails. Feeding the tool a client’s full loan file — income, credit history, family details, prior applications — is more than the task needs. Feeding it the client’s first name, renewal date and mortgage product is enough to produce the same email. The narrower set satisfies the same purpose with materially less personal information exposed to the tool, the vendor operating it, and whatever risk that transfer carries — which is the entire point of the principle, not an afterthought to it.
See also: what is anonymization, what is a data pipeline, what counts as good training data.
Deciding what a live AI system is actually allowed to see, request or log once it is running is an operations control — ai-operations covers how that boundary is enforced day to day.