The question that actually matters isn’t how much data you have. It’s what you’re asking the tool to do with it.
Short answer
There is no fixed number, and no Canadian source publishes one. How much data an AI application needs depends entirely on the task — whether you’re using a pre-trained model as-is, retrieving from your own documents, or fine-tuning something custom, each has a completely different data appetite.
A large share of practical business AI use doesn’t train anything at all — it prompts or retrieves against an existing model, which can need surprisingly little proprietary data to be useful. Fine-tuning a model on your own material is a different exercise with a different data appetite, closer to how a model is actually built from data than to everyday use of a chatbot.
Reached through the OPC’s own footnote on evaluating a tool’s validity for its intended purpose, the U.S. National Institute of Standards and Technology’s AI Risk Management Framework frames this as matching data to context rather than hitting a threshold: a system must be “valid and reliable… throughout the intended lifecycle of the tool and across the variety of circumstances in which they are used”, per the principle the OPC cites it for. That is a mechanism, not a figure — and it is a U.S. framework, cited here only because the OPC itself points to it.
Statistics Canada’s research doesn’t measure a data threshold either, but it does measure what predicts successful adoption: “firms using data analytics are 15.0 percentage points more likely to adopt AI than firms that do not” — well ahead of the other capabilities StatCan tested, including cloud computing and advanced robotics. That points at working data infrastructure, not raw volume, as the thing worth building first. See does AI need clean data to start for the related quality question.
It isn’t data volume, according to Statistics Canada’s own Q2 2026 survey of Canadian businesses: “more than 1 in 10 businesses (13.4%) reported cybersecurity or privacy concerns to be a barrier that limits the use of AI”, with cost as the second leading barrier at 10.6%. A shortage of raw data wasn’t large enough to make the published list of top barriers. If your business is stalled on this question, the more likely blocker is one of those two, not a missing data warehouse.
Matching the data you actually have to the task you want automated is where this question becomes a build decision.