Treadstone Associates
Article · 8 min read

Why AI confidence is not accuracy

Nothing about a generative AI tool’s tone tracks whether its answer is correct. A model that is confidently wrong writes in exactly the same voice as a model that is confidently right, because fluent, assured-sounding language is what the system is built to produce — correctness is a separate property it is not built to guarantee, and Canadian courts have already had to put a name to what happens when someone mistakes one for the other.

Treadstone Associates · Updated 2026

Key takeaways

  • • A generative AI system is built to produce fluent, plausible-sounding text. Nothing in that objective requires the content of the text to be true, so tone carries no information about correctness.
  • • Canadian courts have a specific name for confidently wrong AI output: a “hallucination” — defined in Ontario practice direction as including “non-existent cases, mischaracterizations of case law, and fabricated quotations.”
  • • Canada’s Cyber Centre states the mechanism directly: an AI system’s outputs can be incorrect, might not make sense, might not account for relevant factors, and can be biased — independent of how confidently they are phrased.
  • • The fix is not a better-sounding answer. It is verification against an independent source, which is exactly what tone cannot substitute for.

What a confident tone is actually a property of

A generative AI system produces text by predicting a plausible continuation, one token at a time, trained to sound like the fluent, well-formed writing in its training data. That objective rewards confident, well-structured phrasing whether or not the underlying claim is true — a fabricated case citation reads exactly as smoothly as a real one, because the model is not scoring the citation against a database of real cases while it writes. Confidence of tone and correctness of content are produced by different mechanisms entirely, which is why one gives you no information about the other.

Hallucination is a defined failure mode, not an edge case

Ontario’s Superior Court of Justice has already had to name this in a practice direction, after enough lawyers filed AI-assisted material that turned out to be wrong. Its own language: “it occurs when counsel or litigants carelessly rely on fictitious authorities generated by AI, commonly referred to as ‘hallucinations’. Hallucinations can consist of non-existent cases, mischaracterizations of case law, and fabricated quotations.” The court’s response makes the point about tone directly: it does not matter how convincing the fabricated material read. “It is the responsibility of all counsel and litigants to guarantee accuracy when preparing materials for use in court proceedings, and particularly when using AI, regardless of whether they directly interacted with the technology … The court will not tolerate inadvertence in this regard.” A confidently written, professionally formatted, entirely fabricated case citation is still a fabrication, and the fluency of the writing was never evidence to the contrary.

Canada’s cyber security guidance names the same gap, generalized past courtrooms

The Canadian Centre for Cyber Security makes the identical point for any business use of generative AI, not just legal filings. Its guidance lists what a generative AI tool’s outputs can be, without qualification:

  • “can be incorrect”
  • “might not make sense”
  • “might not take certain factors into account”
  • “can be biased”

and follows it with the operative instruction: “You should always be aware of and validate your sources to verify whether the content being presented is accurate.” Nothing on that list is detectable from how the output sounds. A biased or incorrect answer is not written in a hesitant or unconvincing voice — it is written in the same fluent register as a correct one, which is precisely why the Centre’s instruction is to validate the source, not to trust the delivery.

What actually substitutes for tone as a signal

The United States’ National Institute of Standards and Technology — cited here only because Canada’s own Privacy Commissioner points to it, in a footnote to the necessity principle in its generative-AI guidance, for “more information on validity and reliability in AI systems” — describes what actually has to replace tone as a check: “validity and reliability for deployed AI systems are often assessed by ongoing testing or monitoring that confirms a system is performing as intended,” and that risk management “may need to include human intervention in cases where the AI system cannot detect or correct errors” — because the system checking its own confidence is not a check at all. A model does not have an independent view of whether it is right; a verification step against a source the model did not generate does. Building that verification step into a workflow is the practical answer to a problem tone alone cannot solve.

The same mechanism outside a courtroom

None of this is specific to legal filings — a courtroom is simply where the failure became visible enough to force a written rule. The same mechanism runs through any business use of a generative tool: a fluent answer about which regulation applies, what a supplier’s policy says, or how a competitor prices something carries exactly the same risk of being confidently wrong, for the same reason. The accuracy definition NIST’s AI Risk Management Framework uses — “closeness of results of observations, computations, or estimates to the true values or the values accepted as being true” — is a comparison against an outside, independent standard. Tone has no access to that standard; it is generated from the same process that generated the possibly-wrong content sitting next to it. A business that only reads an AI answer for how sure it sounds is applying a test that could not, even in principle, catch the failure it is trying to catch.

A quick way to test it yourself

Ask an AI tool for a specific, checkable fact it has no way to verify from context — a citation, a figure, a source. Read the answer purely for tone: does it read as confident either way? If a fabricated answer and a correct one would both read equally sure of themselves, that is the whole point made concrete — tone was never carrying the information you needed.

Common questions

If an AI answer sounds confident, does that mean it checked its sources?

No. A generative AI system produces fluent, assured-sounding text regardless of whether the underlying claim is true. Canada’s Cyber Centre states plainly that outputs “can be incorrect” and instructs users to “always be aware of and validate your sources” rather than relying on how the answer reads.

Is a hallucination a rare glitch in an otherwise reliable system?

Ontario’s Superior Court of Justice treats it as common enough to warrant a standing practice direction, requiring every filer to guarantee accuracy “regardless of whether they directly interacted with the technology.” The safer assumption for any unverified AI output is that it may contain a fabrication, not that fabrications are unusual.

How do Canadian courts define an AI hallucination?

The Ontario Superior Court of Justice’s practice direction defines hallucinations as including “non-existent cases, mischaracterizations of case law, and fabricated quotations” — fabricated content presented with the same fluency and formatting as accurate content.

Does a longer, more detailed AI answer carry more weight than a short one?

Not on its own. Length and detail are more of the same signal as tone — a longer answer is not independently verified just because it says more. NIST’s framework treats accuracy as closeness to an outside true value, which detail alone does not establish; a detailed wrong answer is still wrong.

Related: what AI explainability really buys you, and why AI systems degrade over time.

Build a verification step tone alone can never provide.

A short call covers what a workable human-check point looks like for your own AI use.