An AI firm asking to use a library of articles, images or code for training is asking for a licence, and Canadian copyright law already has a complete answer for how a licence like that has to work — it just has nothing to do with the word “AI”.
Key takeaways
The question is really two questions stacked together: do you own this content, and if so, what does it take to actually grant someone else the right to use it. The Copyright Act answers the first with a default rule — the author of a work is its first owner, unless it was made by an employee in the course of employment, in which case the employer is — and the second with a formal requirement that gets skipped more often than it should be.
treadstonelaw.ca walks through how copyright arises automatically the moment original material is created, and who ends up holding it. That matters here because a business cannot licence what it does not own: content built by an independent contractor, for instance, is generally owned by the contractor absent a written assignment, even though the business paid for it. A licensing conversation with an AI firm is the wrong moment to discover that half the archive being offered was actually written by a freelancer who never signed anything over.
Section 13(4) of the Copyright Act is explicit: no assignment or grant is valid unless it is in writing signed by the owner of the right. A licence granting an AI firm the right to train on a content library is exactly the kind of interest this provision covers, and an informal understanding — an email thread, a verbal deal struck at a conference — is not enough to make it enforceable. If the arrangement is ever disputed, the party relying on an unwritten permission has nothing to point to.
A software or SaaS-style licence agreement is the closest existing template for this, and data-ownership clauses in SaaS agreements covers the same questions a content-for-training deal needs answered: is the grant exclusive or can the same material be licensed to other AI firms at the same time; does the licence cover training only, or also the right to reproduce excerpts of the licensed content in generated output; is it perpetual or does it run for a fixed term; and does the fee structure track a one-time payment, a royalty, or a recurring rate tied to continued use.
An AI business that later gets bought or sold is diligenced on exactly this question. deavo.ai’s own Canadian AI-deal research names data and model provenance as the sector’s #1 diligence snag, and the single biggest snag buyers flag in that process is training-data provenance: being able to show, deal by deal, where every training input came from and under what rights. A licensor who signs a clean, written, properly scoped grant today is doing the AI firm’s future diligence file a favour — and protecting its own position if that firm is later challenged over what it trained on.
A software licence can usually be terminated cleanly: access is revoked, the relationship ends, both sides move on. A content-for-training licence does not behave the same way once the licensed material has actually been used to train a model, because the influence of that material becomes baked into the model’s weights rather than sitting in a file that can be deleted on request. A termination clause can stop future use of newly licensed material and can require deletion of the raw content itself, but it cannot generally reach back and un-train a model that has already been built on an earlier batch. Anyone licensing content for training should treat the term and termination clauses as governing what happens next, not as an undo button for what has already happened.
An archive rarely consists of only originally created material — back issues of a publication routinely include licensed photography, syndicated wire copy, or quotations cleared under terms that never contemplated AI training. the same first-ownership rule applies as much here as it does to a single article: a business can only licence the rights it actually holds, and it cannot grant an AI firm anything broader than what its own underlying licences from photographers, wire services or contributors permit. Auditing the archive for exactly this kind of third-party material, before quoting a price, avoids handing over rights that were never the licensor’s to give.
A trade publisher is approached to licence its archive of back issues for AI training. Before signing anything, the checklist is: confirm which articles were staff-written (owned by the publisher under the employment default) versus freelance-written (owned by the freelancer unless assigned); draft a written licence naming exactly what is granted — training use, non-exclusive, three-year term, no right to reproduce full articles in output; and keep a record of which pieces were included, since that list is what the publisher will be asked to reproduce if the deal is ever diligenced or disputed years later. A termination clause is drafted to stop the raw archive being used for any further training after the deal ends — it is not, and cannot be, a promise that a model already trained on the first three years of licensed issues will forget them.
Not strictly by law, but the two things that actually protect you — confirming you own what you are licensing and getting the grant in writing with clear scope — are exactly the areas where a templated agreement written for a different deal tends to miss something specific to an AI-training use.
Only if the freelancer assigned or licensed it to you in writing; by default an independent contractor retains copyright in what they create, unlike an employee, so paying for the work is not the same as owning it.
No — that is the point of choosing non-exclusive terms. An exclusive licence would bar the owner from granting the same rights to a second AI firm, so the exclusivity term is one of the first things to decide before signing.
You can cancel the licence itself and stop any further use of the raw material, but cancellation does not retroactively remove influence the already-licensed content has had on a model already trained — which is exactly why the scope and term of the grant matter more here than in a typical software licence.
Related: whether training on content without a licence infringes copyright and what AI vendor terms say about who owns the output.
Ownership and licensing terms are part of how a custom build gets specified, not an afterthought once it ships.