Treadstone Associates
Article · 7 min read

Does AI understand what it says?

A generative model can write a legal-sounding paragraph without knowing what a contract is. Whether that distinction matters depends entirely on what you are about to do with its answer.

Treadstone Associates · Updated 2026

Key takeaways

  • • Canada's own government definition of generative AI describes it as modelling “features from large datasets,” not as reasoning about what those features mean — a description of pattern-fitting, not comprehension.
  • • Fluent, grammatically correct, confidently-worded output is not evidence of understanding. It is evidence of a well-trained pattern; the two produce identical prose and very different reliability.
  • • The U.S. standards body NIST's own definition of validity requires “objective evidence” that a system's output actually meets the requirements of its use — a test the model itself has no way to apply to its own output.
  • • The practical consequence: treat a generative answer as a draft produced by pattern-matching over data, not as a considered judgment, and route it through the same review a draft from an unfamiliar source would get.

What the question is actually asking

“Does it understand?” sounds philosophical, but it has a very practical version: when a model produces a confident, well-formed answer, is there a process behind it that resembles checking whether the answer is true — or is fluency the whole of what happened? The honest answer, grounded in how these systems are described by the bodies that regulate and study them rather than by marketing copy, is closer to the second.

What Canada's own guidance says the mechanism is

The Canadian Centre for Cyber Security's guidance on generative AI gives a plain, official definition, and the wording is precise about what is happening under the hood: “Generative AI is a type of AI that generates new content by modelling features from large datasets that were fed into the model.” The same guidance elsewhere describes the mechanism from the user's side: it “uses machine learning to construct responses based on a prompt or query.” Notice the verb throughout is model and construct, not know or reason. That is not an oversight in the drafting — it is a deliberately accurate description of a statistical process: the system has learned which patterns of words tend to follow other patterns of words, at a scale no person could hold in their head, and it produces its response by continuing that pattern in a way that fits the prompt.

The same document distinguishes generative systems from earlier, narrower ones on exactly this axis, in its own words: “While traditional AI systems can recognize patterns or classify existing content, generative AI can create unique content in many forms, including text, image, audio or software code.” Creating fluent new content from a learned pattern is a genuinely different and more powerful capability than classifying existing content — but it is still, at bottom, pattern completion operating at very large scale, not a claim to comprehension of what the pattern describes.

Why fluent is not the same test as true

A model trained on enough well-written, mostly-correct text will produce grammatically sound, confident-sounding prose almost by construction — that is what the training process optimizes for. Nothing about fluent phrasing implies the system has verified the underlying claim, because the system was never asked to verify anything; it was asked to continue a pattern plausibly. A wrong fact and a right fact can be phrased with exactly the same confidence, because confidence of phrasing and correctness of content are produced by different parts of the process, and only one of those parts was optimized for in training.

This is the same gap the standards literature names when it separates validity from fluency. The U.S. National Institute of Standards and Technology's AI Risk Management Framework defines validation as “confirmation, through the provision of objective evidence, that the requirements for a specific intended use or application have been fulfilled,” and separately warns, in its own words, that “Deployment of AI systems which are inaccurate, unreliable, or poorly generalized to data and settings beyond their training creates and increases negative AI risks and reduces trustworthiness.” Objective evidence and generalization beyond training are both things a model cannot supply about itself mid-answer. They have to be supplied afterward, by a person or a separate check, which is exactly the review step every Canadian source in this cluster keeps returning to.

Worked example

An illustrative comparison, not a reported case. Ask a model to summarize a 40-page policy document and it will produce a clean, well-organized summary whether the document is internally consistent or contradicts itself on page 30. The summary reads the same either way, because producing a fluent summary and checking the source document for internal contradiction are two different tasks, and only the first one is what a language model is actually doing when it writes. A person who has actually read the document catches the contradiction; a fluent summary, by itself, does not.

Why this is also a labelling question, not just a quality one

If a system is producing a pattern-continuation rather than a considered judgment, readers arguably need to know which one they are looking at — and Canadian privacy regulators have said so directly. The OPC's generative-AI principles state that organizations should “Ensure that system outputs that could have a significant impact on an individual or group are meaningfully identified as being created by a generative AI tool” (OPC principles, 2025). That requirement only makes sense if fluent, human-sounding output is not, by itself, a reliable signal of who or what produced the judgment behind it — which is the same gap this article has been describing from a different angle. Canada's voluntary federal code for advanced generative systems goes further for one specific case: a manager of a public-facing system should “Ensure that systems that could be mistaken for humans are clearly and prominently identified as AI systems” (ISED voluntary code) — not law, and binding only on its signatories, but a direct acknowledgment that fluency alone can pass for a human, understanding, judgment behind it.

What this changes about how to use it

None of this means generative output is useless — a well-organized first draft, produced from a pattern learned over enormous amounts of text, is genuinely valuable for many tasks. It means the fluency of the output should not be mistaken for a signal of its accuracy. The two are produced by different mechanisms and have to be checked separately, which is the same structural point the Cyber Centre makes when it tells users to “always be aware of and validate” what a generative system produces (ITSAP.00.041). Whether a given imperfect, unverified-by-construction output is still worth having depends on the task — the subject of can AI be wrong and still be useful? For the mechanical reason the pattern-completion process works so well for language and so poorly for arithmetic, see how AI handles language versus numbers.

Common questions

If AI doesn't understand meaning, why does its output usually make sense?

Because “making sense” grammatically and semantically is exactly the pattern the model was trained to continue — it has processed enormous volumes of well-formed, mostly-coherent text and learned which words plausibly follow which. That produces sensible-sounding prose reliably. It does not require, and does not imply, that the system has checked the content against reality.

Is this the same thing as a hallucination?

It's the root cause behind that label, not a synonym for it. A hallucination is what happens when the pattern-completion process produces a fluent, confident statement that happens to be false or fabricated — a predictable outcome of a system optimized to produce plausible continuations rather than verified facts, not a rare malfunction.

Does a bigger or newer model close this gap?

A larger or more recent model can make the pattern-fitting sharper and the fluent output correct more often across a wider range of prompts, but it does not change what kind of process is happening. There is no published Canadian benchmark for how much any given model closes this gap, and the structural point — that fluency is not evidence of verification — holds regardless of model size.

Where this goes next

Deciding how much review a generative output needs before it's acted on is a strategy question, not a technical one — it depends on what the answer feeds into.