Treadstone Associates
Article · 8 min read

Why AI gets simple things wrong

It is one of the more disorienting experiences of using a generative AI tool: it produces something genuinely impressive — a well-structured argument, a fluent summary of a dense topic — and then, in the very next sentence, gets a simple, checkable fact wrong: miscounts a list it just wrote, misstates a date, or confidently states something false with exactly the same tone it used for everything correct around it. That combination is not a contradiction. It is a direct consequence of what these systems are actually optimizing for.

Treadstone Associates · Updated 2026

Key takeaways

  • • Fluency and correctness are two separate properties. A model is built to produce plausible, well-formed language — producing true language is a byproduct of that, not the direct target.
  • • Canada's Cyber Centre states this explicitly in its own guidance: outputs “can be incorrect,” “might not make sense,” and “might not take certain factors into account,” and a user “should always be aware of and validate” them.
  • • Confidence in tone carries no information about correctness — a model has no separate mechanism that lowers its fluency or hedges its phrasing specifically when it is about to be wrong.
  • • Tasks requiring exact tracking — counting, precise dates, letter-by-letter detail — are hit hardest, for the same structural reason arithmetic is: the model predicts a plausible-looking answer rather than executing a check.

Fluency is the target. Truth is a side effect.

A generative model's core mechanism, covered earlier in this series, is predicting the next plausible token given everything so far. That mechanism was shaped, through training, to produce text that reads as coherent, well-formed and contextually appropriate. It was not built with a separate, independent process that checks each claim against a ground truth before releasing it. Correct answers and fluent answers overlap enormously in practice, simply because the training data mostly consists of text where fluent and true were the same thing — but the two are genuinely separate properties, and the model has no internal alarm that goes off when they come apart.

Canada's own guidance states this without hedging

The Canadian Centre for Cyber Security's generative-AI guidance is unusually direct on this point. It instructs readers to “always keep in mind that its outputs: can be incorrect, might not make sense, might not take certain factors into account, can be biased,” and follows with the operative instruction: “You should always be aware of and validate your sources to verify whether the content being presented is accurate.” (Canadian Centre for Cyber Security, ITSAP.00.041) Notice what this guidance does not say: it does not say outputs are usually wrong, or that the tool is untrustworthy. It says the possibility of a wrong output does not announce itself, and validation is therefore the user's job, not something the tone of the answer will do for you.

Why simple, checkable facts are hit especially hard

Counting the items in a list it just generated, tracking exactly how many days sit between two dates, or reporting how many times a specific letter appears in a word are all tasks that require exact, symbolic tracking — the same structural weak point covered in the previous article on arithmetic. A model predicting a plausible next token can produce a fluent sentence describing the count without having actually performed anything resembling a count. This is precisely why these errors so often look bizarre in isolation: the same system gets a genuinely difficult, nuanced question right and a trivial counting question wrong, because the two tasks are not related in difficulty from the system's point of view the way they are from a person's.

Confident tone is not evidence

Perhaps the most important, and most counter-intuitive, structural fact for a business reader: a model's phrasing does not reliably soften, hedge or change tone when it is about to state something false. The same fluent, assured register is available for a well-supported claim and a fabricated one, because tone is a property of language generation, and correctness is not something the generation mechanism is separately tracking as it writes. Treating confident phrasing as a signal of accuracy is one of the most common and costly habits a new user of these tools can form, and unlearning it — judging claims by whether they can be checked, not by how they sound — is the single most useful habit this article can leave you with.

Why the burden sits with the person checking, not the tool warning you

Canada's federal, provincial and territorial privacy commissioners, in their joint accuracy principle for generative AI, place a specific duty on the organizations that build and provide these systems: to “inform organizations using generative AI about any known issues or limitations about the accuracy of generative AI outputs.” (OPC, generative AI principles, Accuracy) That is a disclosure duty on the vendor, not a real-time warning system inside the tool itself — and the distinction matters. Even the technical definition of accuracy that Canadian regulators point to, from the NIST AI Risk Management Framework, describes it as “closeness of results of observations, computations, or estimates to the true values or the values accepted as being true,” measured against a documented test set after the fact, not something the system checks against reality live, mid-answer, as it generates each sentence. (NIST AI RMF, via the OPC's own footnote 16) There is no in-the-moment mechanism inside the generation process itself that flags a specific sentence as the wrong one, which is exactly why the checking has to happen on the reader's side, every time it matters.

A worked example: right where it's hard, wrong where it's easy

A model is asked to summarize the key differences between two complex regulatory regimes, then, in the same response, asked how many distinct points it just listed. It produces a genuinely well-organized, substantively accurate summary of the regulatory differences — a task requiring real synthesis — and then miscounts its own six bullet points as five. The summary and the count were not produced by two different levels of effort; both came from the identical mechanism generating a plausible next token. The regulatory summary happened to align with well-represented patterns in training data. The count required something closer to an exact symbolic operation, which is exactly the category of task, discussed throughout this series, where fluent prediction and correct computation most reliably come apart.

Related: why AI can’t do arithmetic reliably, why AI answers the same question differently, and, on building a verification step into a real workflow, the AI operations hub.

Common questions

Is there a way to tell, from the answer alone, when a model is about to be wrong?

Not reliably, which is the central point of this article. Tone and fluency do not track correctness. The dependable approach is to treat claims that matter as unverified until checked against a source you trust, regardless of how confidently they are phrased.

Does asking the model to “double-check its work” fix this?

It can sometimes surface an error, because re-generating a response is a genuinely different prediction pass and may not repeat the same mistake, but it is not a reliable verification mechanism — it is another round of the same fluent-prediction process, which can just as easily produce a second, differently wrong answer with equal confidence.

Are newer models less prone to this than older ones?

Capability has generally improved over time, but the structural cause described in this article — predicting plausible language rather than verifying claims — is a property of how these systems generate text at all, not a specific limitation of any one generation of models. No source consulted for this series claims that limitation has been eliminated.

Keep going in the Academy

This is one page in a plain-English series on how AI actually works and where it fits in a Canadian business.

Back to the AcademyAI operations