That a system worked when it was tested is a claim about a moment, not a permanent property of the system. The world an AI tool operates in keeps moving after launch — the data it sees, the policies it has to reflect, the vendor model running underneath it — while the system itself, left unchecked, does not move with it. That gap is not a malfunction. It is the default outcome of doing nothing.
Key takeaways
Validating an AI system means confirming it performed correctly against a specific set of inputs, at a specific point in time. The United States’ National Institute of Standards and Technology — cited here because Canada’s own Privacy Commissioner points to its framework, in a footnote to the necessity principle in its generative-AI guidance, for “more information on validity and reliability in AI systems” — states the consequence of that plainly: “validity and reliability for deployed AI systems are often assessed by ongoing testing or monitoring that confirms a system is performing as intended.” Ongoing is the operative word. A single validation at launch answers “did it work then,” not “does it still work now” — and the second question does not stay answered on its own.
A system trained on one snapshot of data keeps encountering a world that has since moved past that snapshot: new products, new regulations, new customer behaviour, new phrasing that did not exist when the training data was collected. None of that requires anyone to do anything wrong — it is what happens by default when a static model meets a world that does not hold still. On top of that structural drift, Canada’s Cyber Centre names an active version of the same failure: “threat actors can inject malicious code into the dataset used to train the generative AI system.” The same guidance adds that doing so “could also increase the potential for large-scale supply-chain attacks.” A poisoned dataset does not announce itself — the system keeps producing fluent output, and nothing about its behaviour flags that the ground it was built on has been tampered with. The Centre lists buggy code as a related risk on the development side: “software developers may inadvertently introduce insecure and buggy code into the development pipeline” as generative tools get folded into how software gets built and maintained, which is its own slow source of degradation in anything the system touches downstream.
The Treasury Board’s Directive on Automated Decision-Making binds federal departments, not private business, but its structure is worth borrowing regardless of scope. It does not treat an Algorithmic Impact Assessment as a one-time approval: departments must commit to “reviewing, approving, and updating the published algorithmic impact assessment on a scheduled basis, including when the functionality or scope of the automated decision system changes,” and the directive itself is “reviewed every two years.” Built into the rule is the same assumption this article is making: a system approved once does not stay approved by default. Something changing about what the system does, or what it is used for, is treated as an automatic trigger for reassessment, not an optional nice-to-have.
NIST’s AI Risk Management Framework builds an entire function around exactly this problem. Its MEASURE function requires that “approaches and metrics for measurement of AI risks … are selected for implementation starting with the most significant AI risks,” and specifically that “the risks or trustworthiness characteristics that will not – or cannot – be measured are properly documented.” (NIST, AI RMF 1.0, MEASURE 1.1) Its MANAGE function then asks the harder question on an ongoing basis: “a determination is made as to whether the AI system achieves its intended purposes and stated objectives and whether its development or deployment should proceed.” (NIST, AI RMF 1.0, MANAGE 1.1) That is a recurring question in the framework’s own design, not a one-time gate passed at launch and forgotten. NIST does not exempt its own framework from that same logic, either: the document is published, by NIST’s own description, as “intended to be a living document”, with a formal review of its content due “no later than 2028”. (NIST, AI RMF 1.0, Update Schedule and Versions) A system that sounds exactly as confident as it did on day one gives no signal, on its own, about whether it has quietly degraded since.
The reason this problem persists is not that businesses ignore it — it is that a degraded system rarely announces itself. It keeps producing output in the same format, at the same speed, in the same confident register it always used, and nothing about that surface changes when the substance underneath has drifted. Ongoing evaluation is what has to stand in for a signal the system itself does not provide, which is the discipline behind the operational side of running an AI tool rather than just building or buying one: someone has to keep checking, on a schedule, because waiting for the system to visibly fail means waiting for the failure to already have happened somewhere that mattered.
A worked example
A business builds a pricing or triage tool around a vendor’s AI model. The vendor updates the model behind the same API six months later — a routine improvement from the vendor’s perspective. Nothing about the business’s own workflow changed, so nobody re-tests it. If the update shifted how the model handles an edge case the business relies on, the business finds out only when the edge case actually occurs — which is exactly the outcome a scheduled review, rather than a launch-day-only validation, exists to catch earlier.
No. NIST’s AI Risk Management Framework treats validity and reliability as things “often assessed by ongoing testing or monitoring,” not a one-time gate. The Treasury Board’s own directive for federal automated decision systems requires scheduled reviews and an automatic review whenever a system’s functionality or scope changes.
Several separate mechanisms, not one: the real-world data the system encounters drifts away from its original training snapshot, an underlying vendor model can be updated without notice, and Canada’s Cyber Centre specifically flags training-data poisoning as an active risk that can be introduced without producing any visible warning sign in the system’s output.
Yes, by the logic the federal directive uses for its own systems — a change to functionality or scope is treated as an automatic trigger for review. A business relying on a third-party model has less visibility into when that happens, which is itself a reason to test periodically rather than assume stability.
Related: why AI confidence is not accuracy, and who is accountable when AI decides.
A short call is enough to map which of your systems have never been re-tested since launch.