No — and this is a different question from whether a project has enough data to start, or whether that data has to be clean.
Short answer
No. Past a certain point, more data does not reliably make an AI system better, and uncurated data can make it actively worse. Canada’s own federal guidance for generative AI developers treats assessing and curating a dataset as a distinct, named safeguard, not an afterthought to collecting more of it.
This is not the same question as whether a project has enough data to get started, or whether that data needs to be spotless — see how much data does AI need and does AI need clean data to start for those. This one is narrower: once you already have data, does simply adding more of it make the result better on its own.
Canada’s Voluntary Code of Conduct for advanced generative AI systems lists, among the measures every developer owes regardless of whether a system is public-facing, an obligation to “assess and curate datasets used for training to manage data quality and potential biases.”
That is a deliberate constraint on what goes into a dataset, not a target for how much goes in. The Code treats the two as separate concerns on purpose, and lists curation as work a developer does regardless of dataset size.
The characteristics document behind the NIST AI Risk Management Framework — a U.S. framework, cited here only because it is the one Canada’s own privacy guidance points to for evaluating a tool’s validity — notes that “deployment of AI systems which are inaccurate, unreliable, or poorly generalized to data and settings beyond their training creates and increases negative AI risks.” (NIST AI RMF, AI Risks and Trustworthiness)
Adding more examples that all look like the data a system already has tends to make it more confident inside that narrow pattern, not more capable outside it. A fraud-screening model trained only on last year’s known fraud patterns, however many more copies of that same pattern it is shown, is no better equipped to catch a genuinely new one it has never seen — more of the same is not the same as more coverage.
See what counts as good training data and where a custom build actually needs it.