Treadstone Associates
Article · 9 min read

Does AI training infringe copyright?

Whether training an AI model on someone else’s copyrighted material infringes copyright is the central unresolved argument in AI law — and in Canada, the argument starts from a statute that was never written with the question in mind.

Treadstone Associates · Updated 2026

Key takeaways

  • • A full-text search of the Copyright Act finds zero occurrences of “text and data” (mining), “machine learning” or “artificial intelligence” anywhere in the statute.
  • • Canada has no text-and-data-mining exception of the kind other jurisdictions have debated — training is analysed, if at all, under the ordinary fair-dealing provision.
  • • Fair dealing in Canada is a closed list of purposes — research, private study, education, parody or satire — that does not include commercial use or AI training as such.
  • • A narrow exception for non-commercial user-generated content exists, but its own conditions — non-commercial purpose, attribution, no substantial market effect — do not fit large-scale commercial model training.

Start with what the statute actually contains, because the honest starting point here is an absence, not a rule. a full-text search of the Copyright Act — 350,000-plus characters of consolidated text, current to mid-2026 — turns up zero occurrences of “text and data”, zero occurrences of “machine learning”, and zero occurrences of “artificial intelligence”, in any form. Parliament has not written a rule for this question yet, in either direction.

No text-and-data-mining exception

Some jurisdictions have adopted an explicit text-and-data-mining, or TDM, exception that lets a party analyse lawfully accessed copyrighted works computationally without infringing. Canada has never adopted one. That silence matters in both directions: it means nobody can point to a TDM carve-out to say training is automatically permitted, and it also means the analysis falls back onto whatever general exceptions already exist in the Act — principally fair dealing.

Fair dealing is a closed list

Section 29 is unambiguous about what fair dealing actually covers: dealing for the purpose of research, private study, education, parody or satire does not infringe copyright. Those five purposes are the entire list. There is no catch-all for “transformative use”, no carve-out for commercial activity generally, and nothing that names AI training as a purpose. A model trained for a commercial product does not obviously fall under any of the five headings on its face, which is precisely why the question stays open rather than resolved.

The narrow exception that does not quite reach training either

The Act does contain a provision for non-commercial user-generated content, and it is worth reading closely because it is the closest thing to a TDM-style carve-out Canada has. It excuses an individual from infringement for building a new work out of an existing one, but only where the use is done solely for non-commercial purposes, the source is credited where reasonable, the individual had reasonable grounds to believe the underlying material was not itself infringing, and the new work does not have a substantial adverse effect on the market for the original. A commercial model built by a company, trained on millions of works with no per-work attribution, does not fit the shape of this exception — it was written for an individual’s single derivative work, not an enterprise training pipeline.

Why this stays a live argument rather than a settled answer

Nothing above says training infringes, and nothing above says it does not. What the statute shows is a gap: a general reproduction right, a short and closed list of fair-dealing purposes, and a narrow non-commercial exception that does not fit the commercial-training fact pattern. Anyone who states flatly that AI training in Canada is either always lawful or always infringing is asserting something the Copyright Act itself does not say. treadstonelaw.ca’s explainer on fair dealing — useful background on what fair dealing does cover, even though it was written before AI training was a live question — is a fair starting point for the general doctrine, not a specific answer to it.

None of this is unique to Canada in the sense that other countries are debating a version of the same question, but the specific statutory answer is Canadian: no TDM exception, a five-item closed list, and one narrow non-commercial carve-out that was written for something else. A business operating here should treat “is AI training fair dealing” as an open legal question resting on that specific gap, not as something a foreign court’s ruling on a different statute has already settled for Canada.

Why this matters most at the diligence stage

Canadian AI-deal research names data and model provenance as the sector’s #1 diligence snag, and training-data provenance is the item buyers flag most often: being able to show what a model was trained on, and under what rights. A business that cannot answer that question — because the training set was assembled without records of source or licence — carries a real, unresolved legal exposure that shows up as a discount or a dealbreaker the moment someone looks closely, regardless of how the courts eventually settle the underlying question.

Two scenarios that land in different places

A small Canadian software firm builds an internal tool that fine-tunes an open model on its own historical support tickets and product documentation — material it wrote and owns outright. There is no third-party copyright question here at all, because there is no third party’s work involved; the analysis above simply does not apply. Contrast that with a firm that scrapes a competitor’s published pricing guides, technical manuals and marketing copy to train a tool meant to compete with that same competitor. The second scenario sits squarely inside the unresolved question this article describes: none of the five fair-dealing purposes obviously covers competitive commercial training, and the non-commercial exception does not apply to a commercial product built by a business rather than an individual’s own non-commercial project. The difference between the two is not the technology — it is whose material went in.

Common questions

Is there a Canadian court ruling on whether AI training infringes copyright?

None that this article relies on. The safe, sourced position is what the statute itself contains — no TDM exception and a closed fair-dealing list — not a prediction of how a court would eventually rule on a fact pattern that has not been tested here.

Does using publicly available content for training make it lawful?

Being publicly accessible and being free to copy are different things under copyright law generally, and the Copyright Act draws no special line for AI training specifically — a work being visible on the open internet does not, on its own, place it outside copyright protection.

Would licensing the training data solve the problem?

Licensing removes the copyright question for the material actually covered by the licence, provided the grant is in writing and properly scoped, which is why some AI firms are moving toward licensed or clearly-permissioned training sets rather than relying on an unresolved fair-dealing argument.

Does it matter if the model only produces new, original output?

The output question and the training-input question are legally separate. Whether generated output itself qualifies for copyright protection is a different analysis, covered separately, from whether the act of training on someone else's protected material without permission was itself an infringing reproduction.

Related: licensing your own content to an AI firm and what the Copyright Act says about authorship.

Training-data provenance belongs in a diligence checklist.

Where a model's training data came from is one of the first questions a diligence process should be built to answer.