Two tools answering the same question differently isn't a malfunction in either one. It's what you'd expect from how each was actually built.
Short answer
Because they were rarely built the same way. Two AI tools are typically trained on different datasets, built on different underlying models, and generate their text one word at a time using their own probability calculations — so a different answer to an identical question is closer to the expected outcome than a sign that one tool is broken.
Generative tools work by “modelling features from large datasets that were fed into the model”, per the Canadian Centre for Cyber Security’s own guidance. Different vendors feed their models different datasets, on different schedules, with different filtering choices — so two tools have effectively read different libraries before you ever typed your question. Since late 2022, several distinct large language models from different companies have reached the market, each shaped by its own training data and design choices, which is the guidance’s own framing for why this landscape has multiple, non-interchangeable tools rather than one standard answer engine.
The NIST AI Risk Management Framework — the US standard Canadian privacy regulators point to when asking businesses to check a tool before relying on it — states plainly: “Evaluate the validity and reliability of the generative AI tool for the intended purpose”. NIST itself defines validation as “confirmation, through the provision of objective evidence, that the requirements for a specific intended use or application have been fulfilled”, and warns: “Deployment of AI systems which are inaccurate, unreliable, or poorly generalized to data and settings beyond their training creates and increases negative AI risks”. That is a system-by-system determination, not a category-wide guarantee — a tool trained mostly on one kind of material can generalize poorly to a question outside it, while another tool, built differently, may not.
The variation isn’t only between tools. Google’s own technical description of how its models generate text explains why a single tool can answer the same prompt two different ways: “Large language models generate text one word (token) at a time. Each word is assigned a probability score”, and the next word is selected from that distribution rather than fixed in advance. Stack all three factors together — different training data, different validated reliability, and probability-based word selection even within one model — and getting the same answer twice from two different tools would be the surprising outcome, not the expected one. For what this means when you’re trying to build something dependable on top of a model, see Treadstone’s Custom AI Solutions hub, or compare two specific tools directly at is Copilot the same as ChatGPT.
Different models behave differently by design. What you build around one has to account for that.