“More data” is not the same question as “good data.” A dataset can be enormous and still be the wrong dataset — unrepresentative, improperly sourced, or quietly encoding a bias nobody intended. What follows is not a checklist unique to AI; it is the set of properties that already existed in data-quality practice, made more consequential because a model repeats what it is shown at scale.
Key takeaways
Before a business asks whether an AI tool is any good, it is worth asking a narrower question first: was the data it learned from any good? The U.S. National Institute of Standards and Technology’s AI Risk Management Framework — the reference NIST document Canada’s own privacy commissioners point to when assessing an AI tool’s reliability — frames this as validity: “confirmation, through the provision of objective evidence, that the requirements for a specific intended use or application have been fulfilled.” NIST AI RMF characteristics Four properties, not one, decide whether a given dataset clears that bar.
NIST’s framework is explicit that “accuracy measurements should always be paired with clearly defined and realistic test sets – that are representative of conditions of expected use – and details about test methodology.” NIST AI RMF characteristics A dataset drawn from one type of client, one region, or one season of business is not automatically wrong — but a tool trained or evaluated on it should not then be trusted on a materially different population without saying so. The framework also warns that “accuracy and robustness… can be in tension with one another,” meaning a dataset optimized to make a system look precise under narrow conditions can make it brittle everywhere else. NIST AI RMF characteristics
NIST’s framework notes plainly that “training data may also be subject to copyright and should follow applicable intellectual property rights laws.” NIST AI RMF characteristics In Canada, that is not a formality. The Copyright Act contains no text-and-data-mining exception and no provision addressing AI training at all — the closest available basis, fair dealing under section 29, covers only “research, private study, education, parody or satire,” a closed list that does not extend to commercial AI training. Copyright Act, s.29 A dataset assembled by scraping someone else’s published work, without a licence and without falling inside that closed list, is not made safe to use by good intentions or a disclaimer. Where a business licenses or purchases a dataset instead, provenance is simpler to establish — but it should still be documented, not assumed.
Canada’s federal, provincial and territorial privacy commissioners’ joint AI principles state that organizations should “use anonymized, synthetic, or de-identified data rather than personal information where the latter is not required to fulfill the identified appropriate purpose(s).” OPC, generative AI principles That is a design instruction, not a cleanup step: a good training or grounding dataset is built to need as little real personal information as the task allows, from the start, rather than scrubbed for compliance once someone raises the question.
The Canadian Centre for Cyber Security’s guidance on generative AI names data poisoning directly among the risks businesses should watch for: “threat actors can inject malicious code into the dataset used to train the generative AI system… this could undermine the accuracy and quality of the generated data,” with the added risk of “large-scale supply-chain attacks.” Cyber Centre, ITSAP.00.041 The same guidance separately flags an ordinary, non-malicious version of the same problem: “most of the training datasets fed into LLMs come from the open Internet… generated content has a fundamental bias in that only limited amounts of the world’s total data are online and available for AI to use.” Cyber Centre, ITSAP.00.041 Bias, in other words, is not a separate ethical add-on to data quality — it is a data-quality defect with the same shape as any other unrepresentative sample.
Canada’s privacy commissioners make a point worth sitting with: developers should ensure personal information used to train a model “is as accurate as necessary for the purposes,” and add that “the introduction of ‘inaccuracies’ by modifying a dataset to address a known bias (such as by enhancing it with synthetic data) may be preferable to use of the original ‘accurate’ dataset.” OPC, generative AI principles In plain terms: a dataset that faithfully reflects a biased history is not thereby a good training set. Sometimes the more defensible choice is to deliberately rebalance it, even at the cost of literal fidelity to the original records.
A Canadian professional-services firm is building an internal assistant grounded in its own client files. Run against the four properties above: are the files representative of the full client base, or mostly recent, high-value files that skew the tool’s sense of a “typical” matter? Does the firm actually hold the rights to the third-party documents mixed into those files — contracts drafted by outside counsel, licensed industry reports — or only a right to use them for the original matter? Has client personal information been minimized to what the assistant’s task genuinely requires, rather than left in wholesale because deleting it looked like extra work? And has anyone checked whether the file set over-represents one type of client, in a way that would make the assistant quietly worse for every other kind?
Related: where AI training data comes from, why data quality decides AI quality, and structured vs unstructured data, explained.
No. NIST’s framework treats volume as one factor among several — representativeness, provenance and freedom from bias matter independently of size, and a large but skewed dataset can make a tool confidently wrong rather than merely imprecise.
Not on the basis that it is “publicly available.” Canada’s Copyright Act has no text-and-data-mining exception, and fair dealing under section 29 is limited to research, private study, education, parody or satire — a closed list that does not include commercial AI training.
No. The Copyright Act contains no provision mentioning artificial intelligence, computer-generated works, or text-and-data mining at all. Training on third-party content is assessed under the ordinary fair-dealing rule, which was not written with AI in mind.
A short call is enough to walk through representativeness, provenance and privacy exposure before you commit to a build.