It is one of the most reliable ways to catch a generative AI system making a confident, wrong claim: give it an arithmetic problem with more than a couple of digits. The same system that can draft a nuanced email or summarize a dense report will, with real frequency, get a multiplication or a running total wrong — not because arithmetic is conceptually hard, but because of exactly how these systems produce an answer in the first place.
Key takeaways
Canada's Cyber Centre notes that large language models are “a subset of generative AI that has seen significant improvement in recent years,” and that “to create content, LLMs are provided a set of parameters (for example, a query or prompt).” (Canadian Centre for Cyber Security, ITSAP.00.041) That improvement has been real and rapid on language tasks specifically — the kind of pattern-rich, loosely-defined problems the underlying mechanism is well suited to. It has not changed what kind of operation the mechanism performs, which is the point this article works through: rapid progress on fluent language does not, by itself, imply matching progress on exact symbolic computation, because the two are not the same kind of task for a next-token predictor.
The earlier articles in this series covered the core mechanism: a model predicts its next token based on a learned probability distribution over what a plausible continuation looks like, given everything so far. Ask it to write a sentence, and that is exactly the right kind of task — plausible, fluent language is what the mechanism is built to produce. Ask it to multiply a four-digit number by a three-digit number, and the same mechanism is now being asked to predict what a correct-looking digit sequence would be, based on patterns seen during training, rather than to execute the actual multiplication algorithm a calculator or a spreadsheet runs. Those are not the same operation, even though both produce a string of digits as output.
Tokenization — the step, covered earlier in this series, that breaks text into pieces before a model does anything else with it — treats digits mechanically, the same way it treats letters. A long number can be split into token chunks that do not line up with place value at all: there is no guarantee that “thousands,” “hundreds” and “tens” each land inside their own clean token boundary. Prose tolerates this kind of chunking without any loss of meaning — a word split into two token fragments still reassembles into the same word. Arithmetic does not tolerate it in the same way, because exact computation depends on tracking carries and place value precisely across every digit, and the model has no built-in mechanism forcing it to track that the way an algorithm does. It is predicting a plausible answer's shape, digit by digit, the same way it predicts a plausible next word.
A model asked to add two two-digit numbers will very often get it right, simply because short, common arithmetic patterns appear frequently enough in training data that the plausible-looking answer and the actually-correct answer converge most of the time. Stretch the problem to six-digit multiplication or a running total across a dozen line items, and the number of opportunities for the predicted digit sequence to drift from the true computed value grows with every additional digit. This is the opposite of how a calculator behaves — a calculator's reliability does not degrade as the numbers get longer, because it is executing an algorithm, not predicting a pattern.
The structural fix in production use is not to make the language-prediction mechanism itself better at arithmetic — it is to stop asking it to do the arithmetic at all. Anthropic's own developer documentation describes exactly this capability, calling it tool use: “Tool use (also called function calling) lets Claude call functions that you define or that Anthropic provides. Claude determines when to call a tool based on the user's request and the tool's description. It then returns a structured call that your application executes… or that Anthropic executes.” (Anthropic, tool use documentation) In practice, this means a well-built product does not ask the underlying model to compute 4,827 × 316 by predicting a plausible-looking digit string. It recognizes the request as a calculation, hands the actual numbers to a genuine calculator function, and returns the calculator's exact result. The model's job in that flow is deciding that a calculation is needed and phrasing the answer, not performing the arithmetic itself.
A natural assumption is that showing a model more worked arithmetic examples during training would close the gap the way more examples improve it at, say, summarizing documents. This helps at the margins for short, common patterns, but it does not remove the structural mismatch: the model is still predicting a plausible next token rather than executing an algorithm with guaranteed carries and exact place-value tracking. No realistic amount of additional training data changes what kind of operation next-token prediction is. This is precisely why the practical fix that has actually shipped in real products is not more training on arithmetic — it is routing the arithmetic to a component built to compute exactly, rather than to predict plausibly. The NIST AI Risk Management Framework, cited here through the Office of the Privacy Commissioner of Canada's own footnote on evaluating an AI tool's validity and reliability, defines accuracy as “closeness of results of observations, computations, or estimates to the true values or the values accepted as being true” — a definition that, for arithmetic specifically, has an exact right answer to be close to or not, unlike an open-ended writing task where several answers can be equally valid. (NIST AI RMF, via the OPC's own footnote 16) (OPC, generative AI principles)
Ask “what is 4,827 times 316?” of a plain chat interface with no tool access, and you are asking the language-prediction mechanism to guess a correct-looking six-digit answer directly — it may get it right, and it may not, and there is no way to tell which from the confident tone of the response alone. Ask the identical question inside a product that has wired up a calculator tool, and the system recognizes the arithmetic, calls the tool with the two numbers, and reports back the tool's exact result: 1,525,332. Both products may use the same underlying model. The difference in reliability comes entirely from whether the product built the tool-use step in, which is exactly why “can this AI do math” is the wrong question — the better question is “does this product hand arithmetic off to something that actually calculates.”
Related: why AI gets simple things wrong, what happens when you send a prompt, and, on wiring a model up to tools and existing systems, the AI integration & automation hub.
It usually improves the odds on shorter, more common arithmetic, but the underlying mechanism — predicting a plausible digit sequence rather than executing an algorithm — does not change with scale. A model with tool access reliably outperforms an even larger model without it, on arithmetic specifically.
Test it directly: give it a calculation long enough that a wrong guess is plausible — a five- or six-digit multiplication — and check the result against an actual calculator. If it is consistently exact even on numbers well outside common training patterns, a tool is very likely doing the real computation.
Related, but not identical. Both trace back to the model predicting a plausible continuation rather than verifying a fact, which the next article in this series covers directly. Arithmetic is a specific, sharply-defined case of that broader pattern, and one of the easiest to test for yourself.
This is one page in a plain-English series on how AI actually works and where it fits in a Canadian business.