Nothing about a generative AI tool’s tone tracks whether its answer is correct. A model that is confidently wrong writes in exactly the same voice as a model that is confidently right, because fluent, assured-sounding language is what the system is built to produce — correctness is a separate property it is not built to guarantee, and Canadian courts have already had to put a name to what happens when someone mistakes one for the other.
Key takeaways
A generative AI system produces text by predicting a plausible continuation, one token at a time, trained to sound like the fluent, well-formed writing in its training data. That objective rewards confident, well-structured phrasing whether or not the underlying claim is true — a fabricated case citation reads exactly as smoothly as a real one, because the model is not scoring the citation against a database of real cases while it writes. Confidence of tone and correctness of content are produced by different mechanisms entirely, which is why one gives you no information about the other.
Ontario’s Superior Court of Justice has already had to name this in a practice direction, after enough lawyers filed AI-assisted material that turned out to be wrong. Its own language: “it occurs when counsel or litigants carelessly rely on fictitious authorities generated by AI, commonly referred to as ‘hallucinations’. Hallucinations can consist of non-existent cases, mischaracterizations of case law, and fabricated quotations.” The court’s response makes the point about tone directly: it does not matter how convincing the fabricated material read. “It is the responsibility of all counsel and litigants to guarantee accuracy when preparing materials for use in court proceedings, and particularly when using AI, regardless of whether they directly interacted with the technology … The court will not tolerate inadvertence in this regard.” A confidently written, professionally formatted, entirely fabricated case citation is still a fabrication, and the fluency of the writing was never evidence to the contrary.
The Canadian Centre for Cyber Security makes the identical point for any business use of generative AI, not just legal filings. Its guidance lists what a generative AI tool’s outputs can be, without qualification:
and follows it with the operative instruction: “You should always be aware of and validate your sources to verify whether the content being presented is accurate.” Nothing on that list is detectable from how the output sounds. A biased or incorrect answer is not written in a hesitant or unconvincing voice — it is written in the same fluent register as a correct one, which is precisely why the Centre’s instruction is to validate the source, not to trust the delivery.
The United States’ National Institute of Standards and Technology — cited here only because Canada’s own Privacy Commissioner points to it, in a footnote to the necessity principle in its generative-AI guidance, for “more information on validity and reliability in AI systems” — describes what actually has to replace tone as a check: “validity and reliability for deployed AI systems are often assessed by ongoing testing or monitoring that confirms a system is performing as intended,” and that risk management “may need to include human intervention in cases where the AI system cannot detect or correct errors” — because the system checking its own confidence is not a check at all. A model does not have an independent view of whether it is right; a verification step against a source the model did not generate does. Building that verification step into a workflow is the practical answer to a problem tone alone cannot solve.
None of this is specific to legal filings — a courtroom is simply where the failure became visible enough to force a written rule. The same mechanism runs through any business use of a generative tool: a fluent answer about which regulation applies, what a supplier’s policy says, or how a competitor prices something carries exactly the same risk of being confidently wrong, for the same reason. The accuracy definition NIST’s AI Risk Management Framework uses — “closeness of results of observations, computations, or estimates to the true values or the values accepted as being true” — is a comparison against an outside, independent standard. Tone has no access to that standard; it is generated from the same process that generated the possibly-wrong content sitting next to it. A business that only reads an AI answer for how sure it sounds is applying a test that could not, even in principle, catch the failure it is trying to catch.
A quick way to test it yourself
Ask an AI tool for a specific, checkable fact it has no way to verify from context — a citation, a figure, a source. Read the answer purely for tone: does it read as confident either way? If a fabricated answer and a correct one would both read equally sure of themselves, that is the whole point made concrete — tone was never carrying the information you needed.
No. A generative AI system produces fluent, assured-sounding text regardless of whether the underlying claim is true. Canada’s Cyber Centre states plainly that outputs “can be incorrect” and instructs users to “always be aware of and validate your sources” rather than relying on how the answer reads.
Ontario’s Superior Court of Justice treats it as common enough to warrant a standing practice direction, requiring every filer to guarantee accuracy “regardless of whether they directly interacted with the technology.” The safer assumption for any unverified AI output is that it may contain a fabrication, not that fabrications are unusual.
The Ontario Superior Court of Justice’s practice direction defines hallucinations as including “non-existent cases, mischaracterizations of case law, and fabricated quotations” — fabricated content presented with the same fluency and formatting as accurate content.
Not on its own. Length and detail are more of the same signal as tone — a longer answer is not independently verified just because it says more. NIST’s framework treats accuracy as closeness to an outside true value, which detail alone does not establish; a detailed wrong answer is still wrong.
Related: what AI explainability really buys you, and why AI systems degrade over time.
A short call covers what a workable human-check point looks like for your own AI use.