Clean enough for what you’re asking it to do is a more honest bar than “perfect.”
Short answer
Not spotless, and not garbage either — the honest answer is clean enough for the specific task, and what counts as clean enough is set by how the output will be used, not by an absolute standard that applies to every project the same way.
This isn’t only a technical concern. Canada’s joint privacy principles for generative AI tell developers to evaluate “the training data sets to ensure that they do not replicate, entrench, or amplify historical or present biases – or introduce new biases”, warning that skipping this step may be more likely to result in discriminatory outcomes. Bad data isn’t just a quality problem — it’s a compliance one, because the output of a model trained on it can produce discriminatory results. This isn’t hypothetical liability, either: in Ontario, the Human Rights Code s.5(1) gives “a right to equal treatment with respect to employment without discrimination because of race, ancestry, place of origin, colour, ethnic origin, citizenship, creed, sex...age, record of offences, marital status, family status or disability” — a right that does not stop applying because a hiring or screening decision ran through an AI tool trained on biased data instead of a person. This plain-language rundown of the protected grounds covers what counts.
No Canadian source publishes a data-cleanliness benchmark, and this page doesn’t invent one. The closest usable framework is American: the U.S. National Institute of Standards and Technology’s AI Risk Management Framework, reached through the OPC’s own footnote on evaluating “the validity and reliability of the generative AI tool for the intended purpose”, names “valid and reliable” as one of the core characteristics a trustworthy system needs — judged against the specific job the tool is doing, not against a universal threshold.
Statistics Canada’s research on AI adoption offers something more useful than a cleanliness score: “firms using data analytics are 15.0 percentage points more likely to adopt AI than firms that do not”. That’s a real, sourced signal about what actually correlates with successful adoption in Canadian firms — not raw data volume or perfection, but whether the business already has working analytics processes running on the data it has.
See also how much data does AI need for the volume question this one is often paired with.
Getting data into a state a connected tool can actually use is where this question turns into implementation work.