An AI answer doesn't appear from a single step. It is the end of a pipeline that starts with raw data and passes through several distinct stages before anything reaches a user — and each stage is a place things can go right or wrong.
Key takeaways
It is easy to think of an AI system as a black box: type a question in, get an answer out. What actually sits between the raw data an organization starts with and the answer a user eventually sees is a defined sequence of stages, and naming them precisely matters, because each one is a real point where the final answer's quality is decided — long before anyone types a question.
Canada's federal voluntary code for advanced generative AI systems is unusually explicit about this sequence, because it needs to assign accountability across it. Its own footnote defines “development” as covering “methodology selection, collection and processing of datasets, model building, and testing” — in that order. Before a single parameter is trained, someone has already made two consequential decisions: which general approach to use, and which data to collect and how to process it. Both decisions constrain everything that follows — a model can only learn from the data it is shown, so what gets collected here is effectively the ceiling on what the finished system will ever be able to do.
“Model building” is the stage most people picture when they think of “training AI” — the process, described elsewhere in this hub, of adjusting a model's internal parameters against the collected examples until its output fits the pattern in them. Canada's voluntary code names a specific safeguard that belongs to this stage and not a later one: developers should “Assess and curate datasets used for training to manage data quality and potential biases” (ISED voluntary code) — curation happens here, before the pattern is locked in, because it is far cheaper to fix a data problem before training than to correct its effects afterward.
Testing is named as its own distinct stage, separate from building, and NIST's AI Risk Management Framework gives a useful frame for what testing is actually checking: whether the system's output has the “objective evidence” needed to confirm “the requirements for a specific intended use or application have been fulfilled” (NIST AI RMF, a U.S. framework). The framework's own Core function set treats this as a discipline that spans the whole pipeline rather than a single checkpoint: Map establishes and documents context before anything is built (Map 1, NIST AI RMF Core); Measure tracks what can actually be assessed, and the framework says plainly that “The risks or trustworthiness characteristics that will not – or cannot – be measured are properly documented”; and Manage asks, in its own words, whether “the AI system achieves its intended purposes and stated objectives and whether its development or deployment should proceed” at all (Manage 1.1) — a genuine off-ramp, not a formality, at the end of the build phase.
The pipeline does not end when a model passes testing. Canada's voluntary code separates out a distinct later phase, “Managing the operations,” defined as “putting a system into operation, controlling the parameters of its operation, controlling access, and monitoring its operation” — four further activities, named as their own ongoing obligation rather than folded into “development” (ISED voluntary code, footnote 2). This is the stage that produces the actual answer a user finally sees — and it is also the stage that can go quietly wrong long after the model itself was built correctly, which is the subject of why automations break quietly.
The pipeline, stage by stage:
A business that buys access to someone else's model rather than building its own does not skip these stages — it inherits them, already decided, from the vendor. The data collection, the curation, the testing against documented requirements: all of it happened upstream, out of sight, before the tool ever reached a purchase decision. What the buying business is actually responsible for is the later phase — putting the tool into operation inside its own workflow, controlling who has access to it, and monitoring what it produces once real work is running through it. Confusing the two is a common and consequential mistake: assuming a vendor's testing covers your specific use case, when the vendor's testing was necessarily done against its own documented intended use, not yours.
Each of these stages is a real design decision, not a formality — which is why two organizations using what is nominally “the same technology” can end up with very differently behaved systems. For what actually happens inside stage two, see how a machine learns from examples; for the plain-language version of the whole path without the pipeline detail, see how artificial intelligence works, in plain English.
No — training (or “model building”) is one stage inside a longer pipeline that also includes choosing the data and method beforehand, testing afterward, and putting the system into operation and monitoring it once it's live. Training gets the most attention, but it is not the whole process.
It has a clear build phase, but Canada's own federal guidance treats the operational phase — monitoring, access control, ongoing parameter management — as its own standing obligation rather than a one-time step after launch. In practice, a well-run AI system is never fully “finished” the way a piece of conventional software sometimes is.
It depends on whether the business built the model or bought access to one. A business using a vendor's model inherits the vendor's earlier stages — data collection, model building, testing — and is directly responsible for the later “managing the operations” stage: how it's deployed, who has access, and whether it's monitored once live.
Deciding whether to buy access to a model that has already been through this pipeline, or build one shaped around your own data, is the central build-versus-buy question.