Treadstone Associates
Ask an Expert · 4 min read

Does AI need clean data to start?

Clean enough for what you’re asking it to do is a more honest bar than “perfect.”

Treadstone Associates · Updated 2026

Short answer

Not spotless, and not garbage either — the honest answer is clean enough for the specific task, and what counts as clean enough is set by how the output will be used, not by an absolute standard that applies to every project the same way.

Why regulators care about this at all

This isn’t only a technical concern. Canada’s joint privacy principles for generative AI tell developers to evaluate “the training data sets to ensure that they do not replicate, entrench, or amplify historical or present biases – or introduce new biases”, warning that skipping this step may be more likely to result in discriminatory outcomes. Bad data isn’t just a quality problem — it’s a compliance one, because the output of a model trained on it can produce discriminatory results. This isn’t hypothetical liability, either: in Ontario, the Human Rights Code s.5(1) gives “a right to equal treatment with respect to employment without discrimination because of race, ancestry, place of origin, colour, ethnic origin, citizenship, creed, sex...age, record of offences, marital status, family status or disability” — a right that does not stop applying because a hiring or screening decision ran through an AI tool trained on biased data instead of a person. This plain-language rundown of the protected grounds covers what counts.

The structural answer, since no Canadian number exists

No Canadian source publishes a data-cleanliness benchmark, and this page doesn’t invent one. The closest usable framework is American: the U.S. National Institute of Standards and Technology’s AI Risk Management Framework, reached through the OPC’s own footnote on evaluating “the validity and reliability of the generative AI tool for the intended purpose”, names “valid and reliable” as one of the core characteristics a trustworthy system needs — judged against the specific job the tool is doing, not against a universal threshold.

The one real Canadian data point

Statistics Canada’s research on AI adoption offers something more useful than a cleanliness score: “firms using data analytics are 15.0 percentage points more likely to adopt AI than firms that do not”. That’s a real, sourced signal about what actually correlates with successful adoption in Canadian firms — not raw data volume or perfection, but whether the business already has working analytics processes running on the data it has.

See also how much data does AI need for the volume question this one is often paired with.

Where this goes next

Getting data into a state a connected tool can actually use is where this question turns into implementation work.