Text and data mining, usually shortened to TDM, refers to the automated, computational analysis of large volumes of text or data to extract patterns, relationships or statistics — the kind of bulk processing that sits underneath training an AI model on a body of text. It is a term that matters in a copyright context because several other countries have written a specific legal exception for it, and Canada has not.
A full-text search of the consolidated Copyright Act for “text and data,” “machine learning,” “artificial intelligence” and “computer-generated” returns zero occurrences of each. Canada has no text-and-data-mining exception, and no computer-generated-works provision either. The provision that would have to carry a TDM argument instead, fair dealing, is a closed list. Section 29 reads, in full: “Fair dealing for the purpose of research, private study, education, parody or satire does not infringe copyright”. Those five purposes are the entire list. “Commercial use” is not on it, and neither is “AI training” — the list simply does not contemplate either as a stand-alone purpose.
That absence has a direct practical consequence: mining or training on someone else’s copyright-protected text in Canada cannot rely on a TDM-shaped carve-out the way it might under a regime built with one. Whoever does it has to find another basis — a licence from the rights holder, or an argument that the specific use genuinely falls within one of the five enumerated fair-dealing purposes — and that argument has to be made purpose by purpose, not assumed. Two of the five purposes, criticism/review (s.29.1) and news reporting (s.29.2), additionally require the source and, if given in the source, the author to be mentioned; the plain research/private-study/education/parody/satire ground in section 29 carries no such attribution condition.
A Canadian analytics firm wants to train a search-ranking model on a large corpus of published news articles it does not own. In a jurisdiction with a TDM exception, that training might be lawful without a licence provided defined conditions were met. In Canada, section 29’s list has to do the work instead: unless the firm can fit its bulk processing inside research, private study, education, parody or satire — or negotiate a licence with the publishers — there is no TDM-shaped door to walk through, because no Canadian statute built one.
See also: what are moral rights, does AI training infringe copyright, where AI training data comes from.
Whether a specific training or fine-tuning approach needs a licence before it can touch a given dataset is a build-time question — custom-ai-solutions covers what a build contract should confirm before training starts.