Ordinary software runs a rule someone wrote down. A machine-learning model runs a pattern nobody wrote down — one it built for itself by being shown enough examples of the answer.
Key takeaways
There are two fundamentally different ways to get a computer to produce a correct answer. The first is to write the rule yourself: if the invoice total exceeds a threshold, flag it for approval; if the postal code matches this pattern, assign this region. Every rule-based program built before the current wave of AI works this way — a person thinks through the logic and writes it down, and the computer follows exactly what was written, nothing more and nothing less.
The second way is to skip writing the rule and instead show the computer many worked examples of the answer, and let it work out for itself what pattern connects the inputs to the outputs. That second approach is what “machine learning” means, and it is the mechanism behind every generative AI system in current use. Canada's Cyber Centre puts it plainly in its own guidance on generative AI: it is a technology that “uses machine learning to construct responses based on a prompt or query” and that “generates new content by modelling features from large datasets that were fed into the model.” The examples come first; the pattern is built afterward, by the machine, not the programmer.
Concretely, a model in training starts with output that is close to random and is shown an enormous number of examples of the kind of answer it is meant to produce. After each one, its internal parameters — the numbers that determine how it responds to a given input — are nudged slightly, in the direction that would have made its own output closer to the example it was just shown. Repeated across a training set large enough, that nudging process converges on a set of parameters that reliably produces output resembling the examples, for inputs it has never specifically seen before. Nobody wrote down the rule those parameters encode; it emerged from the adjustment process, which is why it is usually impossible to point at a single line of logic and say “this is why the model answered that way,” the way you could with a rule-based script.
Canada's federal voluntary code for advanced generative AI systems is unusually precise about naming the stages of this process, because it needs to assign responsibility across them. Its own footnote defines development as covering “methodology selection, collection and processing of datasets, model building, and testing” — four distinct stages, each one a real point where the outcome can be shaped or go wrong, well before the model is ever shown to an end user.
Because the model's behaviour is built entirely from its training examples, the single most consequential decision in the whole process is which examples it is shown. A model trained mostly on formal business English will produce formal business English far more reliably than colloquial speech, not because it was told to prefer one over the other, but because that is the pattern its examples actually contained. A gap or a skew in the training examples becomes a gap or a skew in the model's behaviour, silently, because the model has no independent way to notice what it was never shown.
Canada's voluntary code treats this as a named, distinct safeguard rather than a footnote to model-building generally: developers of advanced generative systems should “Assess and curate datasets used for training to manage data quality and potential biases” (ISED voluntary code). Curation is listed as its own obligation precisely because the examples are not a neutral input to an otherwise-controllable process — they are most of the process.
Worked example
An illustrative comparison, not a technical specification. Teach a child what a “dog” is by writing a rule — four legs, fur, a tail, barks — and the rule will misfire on the first three-legged dog or the first fox it meets, because the rule is only as good as the list of features someone thought to write down. Teach the same child by showing them a thousand photos labelled “dog” and a thousand labelled “not a dog,” and they build their own sense of the category from the pattern across the examples — one that usually generalizes better to an unfamiliar case, but is also entirely shaped by what happened to be in those two thousand photos. Machine learning is closer to the second approach, at a scale no classroom could reproduce — and it inherits the same vulnerability: whatever wasn't well represented in the examples is exactly where it is least reliable.
A pattern learned from examples is only as good as those examples' resemblance to the cases the model will actually meet afterward. NIST's AI Risk Management Framework names this directly as a source of risk: systems that are “poorly generalized to data and settings beyond their training” create and increase risk and reduce trustworthiness (NIST AI RMF characteristics, a U.S. framework). This is also the reason the process described here explains language far more comfortably than exact arithmetic, covered next in how AI handles language versus numbers, and how the individual stages named above — data, training, a working model, and an answer — fit together end to end in how AI is built, from data to answer.
No — the difference is who supplies the logic. In programming, a person writes the rule directly. In machine learning, a person supplies examples of correct input-output pairs and the training process adjusts the model's internal parameters until its own output matches the pattern in those examples. The resulting behaviour is discovered through that adjustment process, not authored line by line.
Not in the way a database stores records. Training adjusts the model's internal parameters based on the examples; the examples themselves are not retained individually inside the finished model in a retrievable form. The model retains a compressed, statistical pattern extracted across all of them, not a lookup table of the originals.
In rule-based software, a gap in test coverage means an untested code path might behave unexpectedly, but the rule itself was still written deliberately. In machine learning, the training examples are not a check on the logic — they are the entire source of the logic. A gap in the examples is a gap in what the model ever learned to do at all, not just a gap in what was verified.
When an off-the-shelf model's training examples don't match your business's actual cases, the fix usually runs through building or fine-tuning something rather than adopting a generic tool as-is.