What a model does during training is a technical process. What a business is legally allowed to train it on is a Canadian legal question with a closed list of answers — and “AI training” is not on that list.
Key takeaways
Training is the process that sets a model’s parameters against a body of data before it is ever put to work; retraining, or fine-tuning, updates those parameters later against new data without starting over. The mechanics of both are a technical question with no dedicated Canadian rule. What Canadian law does have a great deal to say about is a narrower, prior question: what a business is legally allowed to train on in the first place.
The Copyright Act’s fair-dealing provision is a closed list, and it does not include AI training. Section 29, in full: “fair dealing for the purpose of research, private study, education, parody or satire does not infringe copyright.” (Copyright Act, s.29) That is the entire list of permitted purposes — commercial use is not on it, and neither is training a model, however the training is described. A general treadstonelaw.ca explainer on how fair dealing works for Canadian businesses covers the same enumerated-purposes structure for ordinary business use of copyrighted material; the AI-training question sits inside that same closed list, it does not sit outside it. (treadstonelaw.ca, fair dealing in Canadian copyright law)
Canada’s privacy commissioners address the training-data pipeline directly, and their language does not treat public availability as a free pass: they “have recently called on organizations to exercise great caution before scraping” personal information that is merely “publicly accessible,” noting that such information “is still subject to data protection and privacy laws in most jurisdictions,” and that “such scraping is common practice when training generative AI systems.” (OPC, generative AI principles, 2025-05-06) That principle runs on a separate track from copyright — a scraped dataset can raise a privacy problem, a copyright problem, or both at once, and clearing one does not clear the other.
Beyond what a business is allowed to collect, the principles address what training does to the data once collected: developers must evaluate “the training data sets to ensure that they do not replicate, entrench, or amplify historical or present biases – or introduce new biases,” a concern the principles tie specifically to decisions in “health care, employment, education, policing, immigration, criminal justice, housing or access to finance.” (OPC, generative AI principles, 2025-05-06) And the accuracy principle applies to training data specifically, not just to a finished output: personal information used to train a model must be “as accurate as necessary for the purposes.” (OPC, generative AI principles, 2025-05-06)
A model’s parameters are fixed at whatever point its training data was assembled; the world keeps moving, and a model that is never updated gradually answers yesterday’s questions with yesterday’s facts. Retraining or fine-tuning is how an organization brings a model’s behaviour back in line with new data — new products, new terminology, new edge cases the original training data never covered. NIST’s framework treats this as an ongoing governance obligation rather than a one-time event: “a determination is made as to whether the AI system achieves its intended purposes and stated objectives and whether its development or deployment should proceed,” and “legal and regulatory requirements involving AI are understood, managed, and documented” on a continuing basis, not just at launch. (NIST AI RMF Core, via OPC footnote 16)
Fair dealing’s closed list matters when the content being trained on belongs to someone else. It is a different analysis when a business fine-tunes a model on its own historical records — five years of its own customer-support transcripts, for instance — because copyright infringement requires copying another party’s protected work, not using material the business already owns the rights to. The question that actually attaches to a business’s own data is closer to authorship and ownership: who first owned the copyright in material an employee or contractor created. The Copyright Act answers that directly: “the author of a work shall be the first owner of the copyright therein,” except that for work made in the course of employment, “the person by whom the author was employed shall, in the absence of any agreement to the contrary, be the first owner.” (Copyright Act, s.13) A contractor is not an employee for this purpose by default — which is why a written assignment matters for anything a contractor built.
A company wants to fine-tune a model on two sources: five years of its own support-ticket transcripts, and a licensed industry glossary it does not own outright. The first source raises an ownership question, not a fair-dealing one — if the transcripts were written by employees in the course of their employment, the company is the first owner under section 13(3) and can use them freely; if a contractor wrote the underlying support scripts without an assignment clause, the company may not actually own what it is training on. The second source raises the opposite question — fair dealing’s closed list of research, private study, education, parody or satire does not include “training a commercial model,” so using the glossary beyond whatever the licence actually permits is a straightforward infringement risk, regardless of how the use is described internally.
Related: how AI uses a knowledge base, how AI handles messy business data, and the Custom AI Solutions hub
Public availability is not itself a legal green light. Canada's privacy commissioners have specifically flagged scraping publicly accessible personal information as something requiring great caution, and separately, the Copyright Act's fair dealing provision does not include AI training among its permitted purposes.
Not typically a fair-dealing question, because that provision addresses using someone else's protected work. The real question is ownership: whether the material was created by employees in the course of employment (the company owns it by default) or by contractors (it may not, absent a written assignment).
Training sets a model's parameters against an initial body of data; retraining or fine-tuning updates those parameters later against new data, without starting from scratch. NIST's framework treats this as an ongoing governance step, not a one-time event.
Ownership and licensing questions belong at the start of a build, not after the data is already in use.