Treadstone Associates
Article · 8 min read

What “personal information” means for AI

Federal privacy law defines personal information in eleven words. Applying that definition to an AI tool raises a question most businesses haven't had to ask before: does a prediction the model makes ABOUT someone count the same way their name does?

Treadstone Associates · Updated 2026

Key takeaways

  • • PIPEDA's definition is short and broad: “personal information means information about an identifiable individual.” It is not limited to names, addresses or account numbers.
  • • The Privacy Commissioner's own generative-AI guidance states directly that an AI system's output about a person — not just what was fed into it — is itself a collection of personal information requiring legal authority.
  • • Which statute defines “personal information” depends on where the organization operates: PIPEDA federally, or a province's own private-sector Act in Alberta, British Columbia or Quebec.
  • • Free-text prompts, voice recordings, and data that looks anonymized but can be traced back to a person all commonly get missed — not because the definition is unclear, but because an AI interface doesn't look like the structured forms people are used to screening for personal information.

The definition itself

The Personal Information Protection and Electronic Documents Act defines the term at section 2(1): “personal information means information about an identifiable individual” (PIPEDA s.2(1)). The Privacy Commissioner's own summary of Canadian privacy law repeats the same wording verbatim (OPC, Summary of privacy laws in Canada). There is no list of qualifying categories, no size threshold, and no requirement that the information be sensitive, financial, or health-related. If a piece of information can be connected back to a specific, identifiable person, it is personal information under the Act — full stop.

That breadth is easy to underestimate with a traditional customer-data system, where personal information mostly lives in predictable fields: name, email, phone, address. An AI tool breaks that pattern, because it processes free-form text, audio and images — formats that were never designed to separate personal information from everything else.

An AI system's own output can be personal information

This is the point most businesses miss. The Privacy Commissioner's generative-AI principles state it directly, under Principle 1, Legal Authority and Consent: organizations should “be mindful that the inference of information about an identifiable individual (such as outputs about a person from a generative AI system) will be considered a collection of personal information, and as such would require legal authority” (OPC, Principles for responsible, trustworthy and privacy-protective generative AI). In plain terms: if an AI tool reads a customer's file and produces a summary, a risk score, a sentiment rating or any other characterization of that person, the tool has just created a new piece of personal information — one that did not exist in that form before, and one that still needs a lawful basis to exist.

That reframes a common assumption. A business might believe it only has to think about privacy law at the point of collecting raw data — a form field, an uploaded document. The OPC's principle says the output side counts too, which matters directly for what a privacy breach looks like with AI: the thing that gets exposed in an incident is sometimes the model's inference about a person, not only the source data it was built from.

Which statute's definition actually applies

PIPEDA is the federal default, but three provinces run their own general private-sector privacy laws instead. The Privacy Commissioner is explicit: “Unless the personal information crosses provincial or national borders, PIPEDA does not apply to organizations that operate entirely within: Alberta, British Columbia, Quebec. These three provinces have general private-sector laws that have been deemed substantially similar to PIPEDA” (OPC, Summary of privacy laws in Canada). Alberta's and BC's Personal Information Protection Acts, and Quebec's Law 25, each define and apply the concept in their own text — they are not simply PIPEDA under a different name, even though the underlying idea (information about an identifiable person) is the same starting point in all four regimes. A business should confirm which statute actually governs it before assuming PIPEDA's specific mechanics apply.

What commonly gets missed with an AI tool specifically

  • Free-text prompts and chat transcripts. Canada's Cyber Centre flags this directly as a risk: “Users may unknowingly provide sensitive corporate data or personally identifiable information (PII) in their AI queries and prompts” (ITSAP.00.041). Nobody screens a prompt box the way they screen a form field, which is exactly why personal information ends up in it.
  • Voice and image inputs. A recorded call or an uploaded photo can identify a person just as directly as a name on a form, even though neither looks like a database record.
  • “Anonymized” data that isn't, once it's re-identifiable. The Privacy Commissioner's own guidance names model inversion — extracting personal information back out of a trained model — as a live risk, which means data that looks scrubbed at the input stage is not automatically safe at the output stage.
  • Business information that is really about a sole proprietor or a named employee. “Company data” fed into an AI tool often contains information about identifiable people once you look past the corporate label on the file.

What doesn't count

The definition has a real edge, not just a broad middle. Purely aggregate or statistical information — a hypothetical figure like the share of callers in a given month who asked about shipping delays — is not personal information, because no individual is identifiable from it. Information about a corporation as a legal entity, rather than about a person, is generally outside the definition too, unless the information is granular enough to point back at an identifiable individual within or connected to that corporation — a named sole proprietor, a specific employee's file, a particular signatory. The practical test an organization has to apply, repeatedly, is narrower than “is this business data” or “is this sensitive”: it is simply whether a specific, identifiable person can be connected to the information, directly or by combination with something else the organization holds.

A worked example

A business uses an AI tool to draft performance summaries from raw manager notes about individual employees. The raw notes are personal information under PIPEDA's definition — they are, after all, information about an identifiable individual. Less obviously, the AI-generated summary is also personal information in its own right, under the OPC's inference principle above, even though it is a new document the model produced rather than something a person typed. That means the summary itself needs the same lawful basis, the same access rights for the employee it describes, and the same protection as the notes it was built from — it doesn't get a lighter standard just because software wrote it.

Ontario businesses working through what specifically counts as personal information day to day, outside the AI context, can start from the general PIPEDA answer (treadstonelaw.ca, what counts as personal information under PIPEDA); it is written for the general question, and the AI-specific wrinkles above sit on top of that same statutory definition.

Common questions

Is an email address alone personal information?

Yes, if it can identify a specific individual — which most email addresses can, directly or by combination with other information the organization holds. PIPEDA's test is whether the information is about an identifiable individual, not whether it looks sensitive.

If we strip names out before feeding data to an AI tool, is it safe to use freely?

Not automatically. Information can remain personal information if a person is still identifiable from what's left, and the Privacy Commissioner names model inversion — extracting information back out of a trained model — as a specific risk to consider, not a theoretical one.

Does this definition change if the AI tool is used only internally, never shown to a customer?

No. PIPEDA's definition and its obligations attach to the information itself, not to who eventually sees the output. Internal-only use can affect the probability-of-misuse analysis in a breach, but it doesn't remove the information from the definition.

Where this goes next

Knowing exactly what counts as personal information is the first question in any review of how an AI tool or an AI-touched target actually handles data.