Treadstone Associates
Article · 8 min read

How AI handles messy business data

There's no Canadian standard for how clean a data set has to be before an AI tool can be pointed at it. There is an existing accuracy obligation on the data itself — and a distinct risk in how AI fails when that obligation isn't met.

Treadstone Associates · Updated 2026

Key takeaways

  • • PIPEDA's accuracy principle requires personal information to be “as accurate, complete, and up-to-date as is necessary for the purposes for which it is to be used” — a purpose-relative test, not a fixed bar.
  • • Canada's privacy commissioners apply the identical standard to AI training data specifically: information used to train a model must be “as accurate as necessary for the purposes.”
  • • Messy data creates a distinctive AI risk: a human catches an obvious duplicate on sight; an AI system can silently merge or split records wrong, with a confident-looking result either way.
  • • NIST's framework requires documenting which risks “will not – or cannot – be measured” — naming what you haven't resolved, rather than assuming the tool absorbs it.
  • • Accountability for the underlying data stays with the organization under PIPEDA, regardless of which tool processed it.

No Canadian regulator publishes a threshold for how “clean” a data set has to be before an AI system can safely be pointed at it. What exists instead is an obligation that already applied before any AI tool touched the data, and does not go away because a tool is now involved: PIPEDA’s accuracy principle. “Personal information shall be as accurate, complete, and up-to-date as is necessary for the purposes for which it is to be used,” and “information shall be sufficiently accurate, complete, and up-to-date to minimize the possibility that inappropriate information may be used to make a decision about the individual.” (PIPEDA, Schedule 1, Principle 6 (Accuracy))

The obligation is purpose-relative, not a fixed bar

Notice what the principle does not say: it does not set a duplication rate, an error tolerance, or a completeness percentage. It ties accuracy to “the purposes for which it is to be used” — a mailing list tolerates a stale address in a way a credit decision cannot. That means the same messy data set can be adequate for one AI application and inadequate for another, and the test travels with the use, not with the data alone.

Canada’s privacy principles restate the same test for AI training specifically

Canada’s federal, provincial and territorial privacy commissioners apply the identical logic to AI training data directly: developers should “ensure that any personal information used to train their generative AI models is as accurate as necessary for the purposes.” (OPC, generative AI principles, 2025-05-06) The wording is deliberately the same as PIPEDA’s general accuracy principle — this is not a new, AI-specific standard, it is the existing one applied to a new kind of use.

Where duplication and inconsistency create a distinct problem

A human reviewer skimming a spreadsheet usually notices an obvious duplicate — two rows for “J. Singh” and “Jasvir Singh” at the same address stand out because a person reading both rows can tell they probably describe one customer. An AI system processing the same records at scale does not necessarily get that judgment call right, and it does not necessarily get it visibly wrong either — it can silently merge two different customers who share a name, or silently treat one customer as two, producing an answer that looks confident and complete either way. That is the structural risk messy data creates for an AI system specifically: not that it fails obviously, but that its failure mode is quieter than a spreadsheet error a human would catch on sight.

Documenting what you cannot measure

NIST’s AI Risk Management Framework — a United States framework, cited here only through the chain Canada’s privacy commissioners themselves point to — puts a version of this discipline into a formal requirement: “the risks or trustworthiness characteristics that will not – or cannot – be measured are properly documented.” (NIST AI RMF Core, via OPC footnote 16) Applied to messy data, that means naming, in writing, which inconsistencies in the source records the organization has decided not to resolve before deployment — rather than assuming a more capable AI tool will simply absorb them without consequence.

Accountability does not transfer to the tool

Whatever an AI vendor’s marketing implies about cleaning up data automatically, the organization that owns the underlying personal information remains accountable for it under PIPEDA’s accountability principle, whether the tool is a simple deduplication script or a sophisticated model — a designated individual inside the organization is accountable, and that does not change because a vendor’s tool did the processing. (PIPEDA, Schedule 1, Principle 1 (Accountability))

Necessity and proportionality apply before the accuracy question does

Canada’s generative-AI principles frame a prior question that is easy to skip past on the way to a data-cleanup project: is the AI system “necessary and proportionate” for the task at all? “The tool should be more than simply potentially useful. This consideration should be evidence-based and establish that the tool is both necessary and likely to be effective in achieving the specified purpose.” (OPC, generative AI principles, 2025-05-06) Applied to messy data specifically, that principle argues for asking whether a simpler, rules-based deduplication pass would achieve the same result before reaching for an AI system to interpret inconsistent records — not every messy-data problem is an AI problem.

A worked scenario

A retailer’s customer database has accumulated a decade of manual entry errors: some customers appear under two spellings of their name, some records have an outdated address that was never updated after a move, and a batch of records imported from an acquired competitor use a different postal-code format entirely. Before pointing an AI-based customer-service tool at this data, the accuracy principle asks a specific question — is this data “as accurate, complete and up-to-date as is necessary” for what the tool will do with it? If the tool only drafts marketing copy, a stale address may not matter. If it is deciding who qualifies for a loyalty-tier discount tied to purchase history under a specific name, the duplicate records could produce a wrong decision about a real customer — and the retailer, not the AI vendor, remains accountable for that outcome.

Related: how AI models are trained and retrained, how AI uses a knowledge base, and the AI Integration & Automation hub

Common questions

Is there a Canadian standard for how clean data must be before using AI on it?

No numeric standard exists. PIPEDA's accuracy principle instead ties the requirement to the purpose the data will be used for — the same data set can be adequate for one AI application and inadequate for another.

Does more data fix messy data?

Not by itself. Volume doesn't resolve duplication or inconsistency in the underlying records; it can make the problem harder to spot, because a large data set gives an AI system more opportunity to produce a confident-looking answer built on inconsistent inputs.

Who is responsible if an AI tool acts on inaccurate data?

The organization, not the AI vendor. PIPEDA's accountability principle keeps a designated individual inside the organization responsible for personal information regardless of which tool processed it.

Should every messy-data problem be solved with an AI tool?

Not necessarily. Canada's generative-AI principles frame necessity and proportionality as a prior question — the tool has to be shown as necessary and likely effective for the purpose, not merely potentially useful, which argues for trying simpler fixes first.

Connecting AI to data you already have

What a connector can actually reach, and what it does with what it finds there, matters as much as the tool itself.