Anonymization is the process of altering personal information so thoroughly that it can no longer, even indirectly, be linked back to the individual it was originally about — a higher bar than simply removing a name, and one that only one Canadian jurisdiction currently defines by statute.
Federally, PIPEDA defines the term it actually regulates, personal information, rather than anonymization itself: “personal information means information about an identifiable individual”. Information that has been genuinely anonymized is no longer about an identifiable individual, which is why the federal, provincial and territorial privacy commissioners’ own generative-AI principles list anonymized information as an alternative to personal information rather than as a type of it, advising in full: “Use anonymized, synthetic, or de-identified data rather than personal information where the latter is not required to fulfill the identified appropriate purpose(s)”. PIPEDA itself, though, does not set out how anonymized has to be measured.
Québec does. Section 23 of the province’s Act respecting the protection of personal information in the private sector, added by Law 25, is the only Canadian statutory anonymization standard: “information concerning a natural person is anonymized if it is, at all times, reasonably foreseeable in the circumstances that it irreversibly no longer allows the person to be identified directly or indirectly”, and further requires: “Information anonymized under this Act must be anonymized according to generally accepted best practices and according to the criteria and terms determined by regulation”. That is a statutory test — irreversibility, assessed against reasonably foreseeable circumstances, not just a removed name — that no equivalent federal provision currently states.
A national lender strips names, addresses and account numbers from a loan-performance dataset before handing it to an AI vendor for model training, but leaves in postal code, income band, loan amount and approval date for every record. If those remaining fields are distinctive enough that a sufficiently motivated party could re-identify a specific customer by cross-referencing other available data, the dataset has been de-identified, not anonymized in Québec’s statutory sense — the section 23 test asks whether re-identification is reasonably foreseeable, not whether the obvious identifiers were removed.
See also: what is data minimization, what is data residency, where AI training data comes from.
Confirming how thoroughly a target company’s training or customer data has actually been anonymized, versus merely stripped of obvious names, is examination work — ai-due-diligence covers how that check gets run.