Treadstone Associates
Article · 8 min read

How AI handles language versus numbers

A generative model can draft a persuasive paragraph and misstate a two-line sum in the same breath, because writing well and computing correctly are not variations on one skill for this technology — they are two different kinds of problem.

Treadstone Associates · Updated 2026

Key takeaways

  • • Language tolerates approximation — a near-miss sentence is still useful. Arithmetic does not — a single-digit slip invalidates the entire result, with no partial credit.
  • • A generative model works by predicting a plausible next token from a learned pattern, which suits language's statistical redundancy well and suits exact computation poorly, because computation has no tolerance for a highly plausible near-answer.
  • • StatCan's own national survey shows the gap in real deployment: language-shaped uses like text analytics (34.5%) and natural language processing (27.0%) far outrun exact computation-adjacent categories like decision-making systems (13.7%) among Canadian businesses using AI.
  • • The fix in practice is architectural, not aspirational — route the arithmetic to a calculator or a formula, and let the model handle the language wrapped around it.

Two tasks that look similar and aren't

Writing a sentence and computing a sum both look, from the outside, like “produce the right sequence of symbols.” They are not the same kind of problem for a system that works by learning a pattern from examples, and the difference explains a pattern every user of these tools eventually notices: fluent, persuasive prose next to a confidently wrong arithmetic answer, produced by the exact same system in the same conversation.

Why language is a comfortable fit for pattern completion

A generative model, as Canada's Cyber Centre describes it, works by “modelling features from large datasets” and constructing its response by continuing a learned pattern. Language is unusually forgiving of that approach, because natural language is itself highly redundant and statistically patterned: certain words reliably follow certain other words, certain grammatical structures recur constantly, and there is rarely exactly one correct next word — there is a wide range of equally good ones. A model that has absorbed enough of that redundancy can produce fluent, sensible continuations reliably, because the task itself has a large margin for approximation. A sentence that is 95% aligned with the ideal phrasing still reads as a good sentence.

Why arithmetic is a bad fit for the same mechanism

Exact computation has none of that margin. There is exactly one correct answer to “what is 847 times 63,” and every other number is entirely wrong, not approximately right. A next-token predictor trained on text will have seen enormous amounts of arithmetic-flavoured language — word problems, explanations, worked examples — but it has not been running the underlying algorithm the way a calculator's circuitry does; it has been learning what an answer to this kind of question tends to look like, which is a fundamentally different achievement from actually recomputing the product every time. NIST's AI Risk Management Framework names the underlying issue directly: accuracy is the “closeness of results of observations, computations, or estimates to the true values or the values accepted as being true” (NIST AI RMF, a U.S. framework) — and for a genuinely computational task, “close” is not a passing grade the way it is for a sentence. A near-miss number is simply wrong.

There is a second, more mechanical reason the two tasks diverge. A language model does not read a number the way a person does, digit by digit with place value intact; it breaks input text into chunks called tokens, and a multi-digit number can be split across token boundaries in ways that vary from one number to the next. Two numbers that look equally simple to a person can be tokenized completely differently, which means the pattern the model learned for one does not reliably transfer to the other — a problem that barely exists for ordinary words, because words are exactly the unit the tokenizer and the training data were built around in the first place.

What the real Canadian usage data actually shows

This isn't a theoretical distinction — it shows up directly in how Canadian businesses actually deploy the technology. Statistics Canada's second-quarter-2026 survey of AI use among Canadian businesses found that, among businesses using AI, the most common applications were heavily language-shaped: data analytics at 36.6%, text analytics at 34.5%, virtual agents or chat bots at 28.2%, and natural language processing at 27.0% (Statistics Canada, June 2026). Categories closer to exact computation and structured decision-making trailed well behind: decision-making systems based on AI at 13.7%, deep learning at 13.1%, and robotics process automation at just 5.0%. This is real, sourced Canadian adoption data — not a performance claim about which category works better, but a genuine picture of where businesses have found the technology worth deploying, and it lines up with the structural point above: language-shaped tasks dominate real usage.

Worked example

An illustrative comparison, not a benchmark. Ask a generative model to draft a client email summarizing a file, and to add up three line items inside that same email. The email draft will likely read well on the first try, because a good client email has a wide range of acceptable phrasings and the model has seen an enormous number of similar ones. The sum has exactly one correct value, and the model arrived at whatever number it wrote the same way it wrote the rest of the email — by pattern completion, not by actually carrying the addition — so there is no structural reason to expect it to be right, and every reason to check it before it goes out.

The practical fix is architectural, not a bigger model

The reliable answer is not to hope a future model closes this gap on its own — it is to route the two kinds of task to the tool actually built for each: let the language model draft the prose, and hand the arithmetic to a calculator, a spreadsheet formula, or a piece of code written specifically to compute it, then let the model wrap the verified number back into fluent language. That split is exactly the shape described in how AI differs from ordinary software: deterministic logic for the part that needs to be exactly right, a learned pattern for the part that benefits from fluency, and neither one asked to do the other's job. The same reasoning about why a model produces fluent output without necessarily verifying it underlies does AI understand what it says?

Common questions

Will a newer or larger AI model eventually be reliably good at arithmetic on its own?

Some newer systems pair a language model with an actual calculator or code-execution tool behind the scenes specifically to close this gap, which is really an architectural fix rather than the underlying language model getting better at computing. There is no published Canadian benchmark for how reliable any given approach is, so the safer default for a real workflow remains: verify any number that matters.

Is this the same reason AI struggles with dates and unit conversions?

It's the same underlying mechanism. Dates and unit conversions are exact-computation problems dressed in language, so they inherit the same weakness as arithmetic — a model can produce a fluent, confident-sounding conversion that is simply wrong, because fluency and correctness are produced by different parts of the process.

Does this mean AI is bad at anything involving numbers at all?

No — describing, summarizing or discussing numbers in prose (“sales grew compared to last quarter”) is a language task and AI tends to handle it well. The weakness is specifically in exact computation, not in language that happens to mention numbers.

Where this goes next

Deciding which parts of a workflow need exact, deterministic logic and which parts can use a language model is a design question that belongs at the build stage, not an afterthought.