“Garbage in, garbage out” predates AI by decades, but AI is what makes it expensive to ignore. A model trained or grounded on unrepresentative, inconsistent, or biased data does not produce a slightly worse version of the truth — it produces a confident, fluent answer that is simply wrong, with no obvious sign anything went wrong at all. That is a mechanism, not an opinion.
Key takeaways
NIST’s AI Risk Management Framework defines validity as “confirmation, through the provision of objective evidence, that the requirements for a specific intended use or application have been fulfilled,” and reliability as the “ability of an item to perform as required, without failure, for a given time interval, under given conditions.” NIST AI RMF characteristics Both definitions point the same way: quality is not a property of the data in isolation, it is a property of the data relative to a specific, stated use. The same framework adds a warning that matters here: “accuracy and robustness… can be in tension with one another in AI systems” — a dataset tuned to look precise under one set of conditions can become brittle the moment conditions shift even slightly.
Where the data is personal information, that standard is not just a technical best practice — it is a specific legal duty. PIPEDA’s accuracy principle requires that personal information “be as accurate, complete, and up-to-date as is necessary for the purposes for which it is to be used,” a standard the Act applies to the record itself, independent of whatever a model later infers from it. PIPEDA, Schedule 1, cl.4.6 A dataset that fails that legal standard produces an AI tool built on information the business was never entitled to treat as reliable in the first place.
NIST’s guidance calls for accuracy measurement to “demonstrate external validity (generalizable beyond the training conditions)” and to be “paired with clearly defined and realistic test sets… that are representative of conditions of expected use.” NIST AI RMF characteristics The mechanism is straightforward: a model has no independent way to know when it has left the conditions its data actually covered. It keeps producing answers in the same confident register whether the underlying pattern still holds or quietly stopped applying pages ago — which is precisely why a data-quality defect surfaces as a wrong answer delivered with total fluency, not as an error message.
NIST’s framework treats a data-fed system as resilient only if it can “withstand unexpected adverse events or unexpected changes in their environment or use” and “degrade safely and gracefully when this is necessary” — and names the specific failure modes to watch for: “adversarial examples, data poisoning, and the exfiltration of models, training data, or other intellectual property through AI system endpoints.” NIST AI RMF characteristics, Secure and Resilient A dataset that was representative and accurate when a tool launched does not stay that way automatically. A business’s client mix shifts, a regulation changes, a data source gets discontinued — and unless someone is watching for it, the tool keeps producing answers calibrated to conditions that no longer hold, with nothing in its output to flag that the ground has moved.
Statistics Canada’s analysis of AI adoption and productivity in Canadian firms found that “firms using data analytics are 15.0 percentage points more likely to adopt AI than firms that do not,” identifying data analytics, R&D, cloud computing, advanced robotics and ICT training as strong predictors of AI adoption. Statistics Canada, AI adoption and productivity in Canadian firms Read this carefully: it is evidence that firms with better data practices are more likely to adopt AI at all — not a measurement of how much better their AI performs once they do. No Canadian source publishes the latter, and treating an adoption correlation as a performance benchmark would misstate what StatCan actually found.
Canada’s privacy commissioners frame this the same way structurally: developers should evaluate training data “to ensure that they do not replicate, entrench, or amplify historical or present biases – or introduce new biases,” because failing to do so “may be more likely to result in discriminatory outcomes based on race, gender, sexual orientation, disability, or other protected characteristics.” OPC, generative AI principles That is a data-quality problem with the same shape as an unrepresentative sample skewing a model toward the wrong season or the wrong customer segment — it is simply a more consequential kind of unrepresentativeness, not a different mechanism.
Two versions of the same internal tool are built from the same firm’s client records. One draws on three recent years of consistently entered, single-format records. The other draws on ten years of records spanning two different systems, inconsistent field conventions, and a period where the firm stopped tracking a field it now needs. The first tool’s outputs generalize reasonably within the conditions its data actually covers. The second tool’s outputs will look equally confident — the interface gives no indication that a third of its inputs are inconsistent or missing — but its accuracy on anything touching the gap years or the discontinued field has no real basis, because the underlying data never had a stable answer to give it in the first place. Better prompting cannot fix that; the defect is upstream of anything a prompt can reach. Nor does swapping in a newer, more capable model fix it — a better model reasoning over the same inconsistent ten years will still have nothing reliable to say about the gap years, because the problem was never the model’s capability in the first place.
Related: what counts as good training data, why business data is rarely AI-ready, and structured vs unstructured data, explained.
No. Prompting shapes how a model is asked to use what it already knows; it cannot supply representativeness, provenance, or accuracy the underlying data never had. The defect sits upstream of the prompt.
No such figure exists in any Canadian source. Statistics Canada’s data links data-analytics practices to AI adoption likelihood, not to AI performance — any specific accuracy-impact number circulating for Canada has been invented rather than measured.
Only if the additional data is itself representative and accurate. More of the same skew reinforces the skew rather than correcting it — volume and quality are different variables.
No. A more capable model reasoning over the same unrepresentative or inconsistent data will still be limited by what that data actually contains. Model capability and data quality are separate variables, and only one of them is fixed by a vendor update.
A short call is enough to identify where the representativeness gaps actually sit before you build on top of them.