The failures that keep showing up in AI systems aren’t speculation — Canadian courts have had to write formal rules specifically because of them, and Canada’s Cyber Centre names several more by category.
Key takeaways
The Federal Court’s AI notice gives "hallucination" a formal definition, in a footnote, because the problem is specific and recurring: “facts, citations, and other content generated by AI that are not true, and have been fabricated by AI in response to a prompt or request.” Ontario’s Superior Court, in its Consolidated Civil Provincial Practice Direction, describes what that looks like in practice: “counsel or litigants carelessly rely on fictitious authorities generated by AI, commonly referred to as ‘hallucinations’. Hallucinations can consist of non-existent cases, mischaracterizations of case law, and fabricated quotations.”
The same Ontario practice direction lists what follows when a fabricated authority makes it into a filing undetected: “the court’s powers include, but are not limited to, public reprimand of the counsel or litigant, the imposition of cost orders, adjourning a hearing or dismissing the matter, the initiation of contempt proceedings, and in regards to counsel, referral to the Law Society of Ontario.” It adds a line worth sitting with: “it is the responsibility of all counsel and litigants to guarantee accuracy when preparing materials for use in court proceedings, and particularly when using AI, regardless of whether they directly interacted with the technology…the court will not tolerate inadvertence in this regard.” The failure isn’t treated as the AI’s fault — it’s treated as a failure of the review that should have caught it.
Canada’s Cyber Centre’s generative AI guidance lists eight named risks. Several apply directly to agent-style deployments specifically, not just to chat output: “poisoned datasets,” where a threat actor injects malicious content into training data, which the guidance notes “could also increase the potential for large-scale supply-chain attacks” once that model is deployed widely; “buggy code,” where a developer “may inadvertently introduce insecure and buggy code into the development pipeline” using an AI coding assistant; and biased content, which becomes a live problem the moment a model is making or influencing a decision rather than only drafting text for a person to review. The same guidance names a fourth failure mode, closer to the courts’ concern than a security bug: “content not clearly identified as being AI-generated can result in the spread of misinformation, disinformation and confusion.” Canadian Centre for Cyber Security, ITSAP.00.041 An unlabelled output that looks identical to something a person wrote is exactly what lets a fabricated claim travel unchecked.
Ontario’s Superior Court didn’t stop at describing the failure — its practice direction points to an existing procedural rule as part of the answer. Rule 06.1 of the Rules of Civil Procedure already required that “a factum shall include a statement signed by the party’s lawyer…certifying that the person signing the statement is satisfied as to the authenticity of every authority cited in the factum,” with a specific carve-out: an authority “published on a government website…on the Canadian Legal Information Institute website (CanLII), on a court’s website or by a commercial publisher of court decisions is presumed to be authentic…absent evidence to the contrary.” The practice direction separately notes that the Law Society of Ontario’s own Futures Committee produced a White Paper, in April 2024, specifically addressing how the Rules of Professional Conduct apply when generative AI assists in delivering legal services. None of this is a new AI-specific rule invented from scratch — it is an existing verification mechanism, repurposed and pointed at a new failure mode as soon as that failure mode became common enough to name.
None of the sources above puts a number on how often these failures occur, and there is no Canadian benchmark that does — the honest framing is structural. A system tuned and tested against the typical, well-formed version of its task can behave unpredictably on the unusual case: the malformed document, the edge-case request, the input nobody thought to test with. The NIST framework Canadian privacy regulators point to for evaluating AI reliability puts the underlying tension plainly: “accuracy and robustness…can be in tension with one another” in an AI system — tuning hard for one can cost the other, which is a structural limit, not a bug that better testing simply removes.
An internal business tool, not a courtroom, asks an AI system to draft a summary referencing a prior internal decision by name. The system produces a plausible-sounding reference to a decision that was never actually made — the same failure category the Federal Court and Ontario’s courts formally named for legal citations, playing out in a business memo instead of a factum. The risk isn’t limited to litigation; it’s a property of the underlying technology, wherever a confident-sounding, specific-seeming claim goes unchecked before someone relies on it.
The equivalent of Rule 06.1’s carve-out translates outside a courtroom too: a business can build the same habit into its own review step by treating a citation to an internal document as unverified until someone actually opens that document and confirms it says what the AI system claims — the same underlying logic the courts already apply to case law, of treating a source as reliable only once it is named and checkable, applied instead to an internal record.
Related: for the design habits that catch this before it reaches anyone, see why AI agents need guardrails and what “human in the loop” really means.
No — it varies. The Federal Court requires a formal Declaration when AI-generated content resembles co-authorship. Alberta’s tri-court notice, by contrast, urges caution and verification but does not require disclosure of AI use at all — a genuine, sourced contrast between two Canadian court systems handling the same underlying concern differently.
No. Canada’s Cyber Centre names several others, including poisoned training data, insecure code introduced through an AI coding assistant, and biased content — each a different mechanism, not a variation on the same one.
It reduces them but doesn’t eliminate them structurally — the NIST framework Canadian regulators reference on this point is explicit that accuracy and robustness can trade off against each other, which is a property of how these systems work, not a gap that more testing alone closes.
No — the courts are simply where the failure became visible enough, and consequential enough, to force a written rule. The underlying mechanism — a confident, specific-sounding claim that turns out to be fabricated — shows up anywhere an AI system drafts something a person then relies on without checking it.
Ongoing monitoring and review are what turn a known failure mode into a caught one.