Treadstone Associates
Definition

What is de-identification?

De-identification is the process of removing, masking or altering the details in a piece of information that would otherwise let someone connect it back to a specific, identifiable person, so the remaining data can be used with less risk to that person’s privacy.

Treadstone Associates · Updated 2026

How it’s used in Canada

Canada’s federal privacy statute, PIPEDA, does not use the word de-identification anywhere in its current text. What the statute does define closely is the other side of the same question: a business must treat information as personal information whenever it is about “an identifiable individual — meaning it can reasonably be linked back to a specific person, whether alone or combined with other information a business has or can access”. De-identification is the practical answer to that test: reduce or remove the links back to the person until the information stops meeting it.

Québec goes further than the federal statute and sets an actual bar for when information counts as fully anonymized rather than merely de-identified. Since a 2023 amendment, the province’s privacy regulator states that organizations “doivent être en mesure d’anonymiser ces renseignements selon les meilleures pratiques généralement reconnues et en fonction des critères et des modalités déterminés par règlement du gouvernement” — must be able to anonymize personal information according to generally recognized best practices and criteria set by government regulation — and that, absent that regulation, “les organismes et les entreprises ne peuvent anonymiser des renseignements personnels”, organizations cannot anonymize personal information at all. No equivalent federal bar exists outside Québec.

Worked example

Suppose a lender wants to share three years of mortgage-renewal outcomes with an outside AI vendor to help build a model. Before the file leaves the building, someone replaces each client’s name, address and mortgage number with a random file ID, buckets exact birth dates into five-year ranges, and drops any free-text notes field where a person might have typed a name. Whether that is enough depends on how many other details are left in combination — a file can still be re-identifiable if enough of the remaining fields, taken together, point back to one person.

Related terms

See also: what is training data, what is a dataset.

Where this leads

Keeping data de-identified is not a one-time step — ai-operations covers what has to stay true about a data pipeline for as long as the system built on it keeps running.