Training or fine-tuning a model on a business’s own customer records feels like an internal, already-owned-the-data question. Under Canadian privacy law it is not automatically one. Training is a use of personal information, and PIPEDA has rules for what a business is allowed to use personal information for beyond the reason it was collected.
Key takeaways
Schedule 1’s Identifying Purposes principle requires that “the purposes for which personal information is collected shall be identified by the organization at or before the time the information is collected”, and its Limiting Use, Disclosure, and Retention principle follows directly from that: “Personal information shall not be used or disclosed for purposes other than those for which it was collected, except with the consent of the individual or as required by law.” (PIPEDA, Schedule 1, clauses 4.2 & 4.5) A customer’s address and purchase history were collected to fulfil an order. Feeding that record into a training set to build a recommendation model is a different purpose, and clause 4.5.1 requires an organization “using personal information for a new purpose” to document that purpose — which in practice means going back for consent, not assuming the original transaction covered it.
The same clause continues: “Personal information shall be retained only as long as necessary for the fulfilment of those purposes”, and clause 4.5.3 requires that information “no longer required to fulfil the identified purposes should be destroyed, erased, or made anonymous.” (PIPEDA, Schedule 1, clauses 4.5 & 4.5.3) A business that keeps every historical customer record specifically because it might be useful for a future model has, in effect, created a new retention purpose — and that purpose has to be identified and consented to in its own right, the same as any other new use. Treadstone Law’s summary of the ten PIPEDA principles states this one in a single line worth keeping close at hand: “Limiting use, disclosure, and retention — Use information only for the stated purpose; don’t keep it longer than necessary.” (Treadstone Law, on PIPEDA’s principles generally)
The joint federal-provincial-territorial privacy commissioners’ generative AI principles put an explicit obligation on whoever is training the system: “When developing a generative AI system, this means evaluating the training data sets to ensure that they do not replicate, entrench, or amplify historical or present biases – or introduce new biases”, warning that skipping this step “may be more likely to result in discriminatory outcomes based on race, gender, sexual orientation, disability, or other protected characteristics”, particularly in contexts such as “health care, employment, education, policing, immigration, criminal justice, housing or access to finance.” (OPC, generative AI principles) A customer database is rarely a neutral sample of the population it serves, which is exactly why the principles treat evaluating it as a required step, not a nice-to-have.
The same principles state the preferred alternative plainly, under necessity and proportionality: “Use anonymized, synthetic, or de-identified data rather than personal information where the latter is not required to fulfill the identified appropriate purpose(s).” (OPC, generative AI principles) If a training objective can be met with de-identified records — and for most pattern-recognition or recommendation purposes it can — the purpose-and-consent problem described above shrinks considerably, because de-identified information that cannot be linked back to an individual falls outside PIPEDA’s own definition of “personal information” in the first place.
A separate question sits underneath “can we train a model”: whether an AI vendor is allowed to train its own general-purpose model on the records a business sends it to get a service done. The generative AI principles’ distinction between “Developers and Providers” and “organizations using generative AI” matters directly here — a vendor training its own model on a customer’s uploaded records has stepped into the developer role, and the client organization is still the one accountable under Schedule 1’s clause 4.1.3 for having sent personal information to a processor without “contractual or other means to provide a comparable level of protection.” (PIPEDA, Schedule 1, clause 4.1.3) A vendor contract that is silent on whether uploaded data trains the vendor’s own future models is not a neutral gap — it is the exact term a business needs to pin down before sending customer records to any AI product, because the default answer for a given vendor is never safe to assume.
A subscription-box retailer wants to fine-tune a recommendation model on three years of order history: what each customer bought, when, and what they returned. The original privacy notice at checkout said data would be used “to process your order and for account support.” Training a recommendation engine on that history is a new purpose under clause 4.5, so the retailer has two honest paths: go back to customers with a specific, understandable disclosure and a chance to opt out before training on their records, or strip the training set down to de-identified purchase patterns — product, category and season, with no name, account ID, or address attached — which moves the whole exercise outside personal-information rules altogether. Quietly training on the existing dataset under the original checkout notice is the option that does not hold up against clause 4.5.1’s documentation requirement.
Related: whether PIPEDA applies when a business uses AI, what meaningful consent requires for AI uses, and what determines whether customer data is safe in an AI tool.
PIPEDA’s definition of personal information in section 2(1) is “information about an identifiable individual” — data that has genuinely been de-identified so an individual cannot reasonably be re-identified from it, alone or combined with other available information, falls outside that definition, which is why the commissioners point to it as the preferred route for training.
Consent addresses the purpose-and-use problem, but the generative AI principles’ fairness and necessity requirements still apply on top of it — consent does not excuse skipping an evaluation of whether the training data itself will produce biased or discriminatory outcomes.
Once a customer’s information should have been destroyed or anonymized under clause 4.5.3’s retention rule, it should not still be sitting in a training corpus — a deletion request is exactly the kind of event that should trigger removal from any dataset built from live customer records, not just from the primary system it was collected into.
A buyer evaluating an AI-touching business will ask exactly this question about its training data — Due Diligence covers what gets checked.