A model does not know anything a business told it directly. Whatever it produces is a reflection of whatever it was trained on — and for most large AI systems, that material was assembled long before a Canadian business ever opened an account, from sources it never chose and mostly cannot see.
Key takeaways
“Where did it learn that?” is a fair question to ask about any AI tool, and the honest answer usually has three parts: the open web, licensed or purchased material, and — increasingly, for a business’s own tools — its own records.
The Canadian Centre for Cyber Security’s guidance on generative AI is direct about this: “most of the training datasets fed into LLMs come from the open Internet. As such, generated content has a fundamental bias in that only limited amounts of the world’s total data are online and available for AI to use.” It adds that “generated content may be prejudiced if the training dataset lacks balanced representation of data points.” Cyber Centre, ITSAP.00.041 That is not a defect a business can fix on its side — it describes the foundation model itself, before any business ever touches it.
Some training material is licensed outright rather than scraped. Where it is not, Canadian copyright law does not simply wave the question through because the end use was AI. The Copyright Act has no exception for text-and-data mining and no provision mentioning artificial intelligence at all; the only relevant doctrine, fair dealing under section 29, applies only to “research, private study, education, parody or satire.” Copyright Act, s.29 NIST’s AI Risk Management Framework — the U.S. reference point Canada’s privacy commissioners cite for evaluating an AI tool’s reliability — makes the same point from the technical side: “training data may also be subject to copyright and should follow applicable intellectual property rights laws.” NIST AI RMF characteristics A business that licenses a foundation model is generally relying on the vendor to have sorted this out; a business assembling its own training set from documents it found online cannot assume the same.
Canada’s privacy commissioners’ joint AI principles distinguish two roles: “Developers and Providers” — those who “determine how a generative AI system operates, how it is initially trained and tested, and how it can be used” — and “organizations using generative AI,” who deploy a tool someone else built. OPC, generative AI principles Most Canadian businesses sit in the second category for the foundation model itself. But the same guidance notes an organization “might shift between or play multiple roles at once” — and that shift happens exactly when a business fine-tunes a model, or builds a retrieval system, on its own client records, contracts, or knowledge base. At that point the business has become a developer of its own AI system, for that data, whatever it remains for the underlying model.
Whether a business is a developer or a user, the same principle applies to anything “publicly accessible” that includes personal information: Canada’s privacy commissioners have “called on organizations to exercise great caution before scraping…‘publicly accessible’ personal information, which is still subject to data protection and privacy laws in most jurisdictions.” OPC, generative AI principles Posting something online does not place it outside privacy law, and “such scraping is common practice when training generative AI systems” is offered as a warning in that guidance, not a defence.
Not all training data traces back to a real document or a real person. Synthetic data — records generated to resemble real data statistically without being drawn from any actual individual’s information — is a real and growing fourth source, and Canada’s privacy commissioners treat it as a legitimate design choice, not a workaround: their guidance notes that “the introduction of ‘inaccuracies’ by modifying a dataset to address a known bias (such as by enhancing it with synthetic data) may be preferable to use of the original ‘accurate’ dataset.” OPC, generative AI principles Synthetic data does not solve the copyright question above — it is usually generated from a real dataset that still has to be lawfully obtained in the first place — but it can reduce how much real personal information a business needs to expose once that first dataset exists.
A Canadian firm is choosing among three ways to build an AI assistant. Buying access to a general-purpose foundation model means inheriting whatever the vendor scraped, licensed, or was given — largely someone else’s governance problem, for better and worse. Fine-tuning that model on the firm’s own five years of client correspondence means the firm has just become the developer of record for that training set, with the copyright, consent, and minimization questions above now sitting on its own desk. Building a retrieval system that looks up the firm’s current documents on demand, rather than baking them into the model’s weights, narrows the governance question to the retrieval index itself — smaller, but not zero, since the same personal-information and provenance questions still apply to whatever sits in that index. A fourth option — generating synthetic client scenarios to test the assistant rather than using real files at all — sidesteps the personal-information question for testing purposes specifically, without touching whether the underlying model was trained lawfully in the first place.
Related: what counts as good training data, how AI models are trained and retrained, and what happens to data you paste into AI.
Not necessarily, and Canada’s Copyright Act does not automatically excuse it — there is no text-and-data-mining exception, and fair dealing covers only research, private study, education, parody or satire. Whether a specific use infringes depends on facts most businesses cannot see from the outside.
It depends entirely on the vendor’s own terms — some enterprise agreements explicitly exclude customer inputs from training, others do not. This is a contract question to confirm directly with the vendor, not something general Canadian law answers uniformly.
Canadian privacy regulators have flagged this as a live concern, not a settled right — they caution organizations to exercise great caution before scraping publicly accessible personal information, but there is no established Canadian mechanism specifically for demanding a model be retrained without it.
A short call is enough to work out which governance questions land on your desk and which stay with the vendor.