An ordinary automation tells you when it's broken. An AI-touched one often doesn't — it keeps producing an answer, and the answer looks fine right up until someone checks it against reality.
Key takeaways
Traditional, rule-based automation tends to fail loudly. A script expects an invoice number in a fixed format; feed it something else and it throws an exception, the job queue backs up, and somebody gets paged. The failure is inconvenient, but it is visible — the system stops doing the wrong thing because it stops doing anything at all.
A step built on a generative model rarely behaves that way. Ask it to classify a lead, summarize a document or draft a response, and it will almost always return something — fluent, on-topic, and structurally correct-looking. Canada's Canadian Centre for Cyber Security describes generative AI as a technology that “generates new content by modelling features from large datasets that were fed into the model” rather than one that checks its output against a fixed rule. There is no built-in equivalent of the exception the script throws. The pipeline downstream receives a well-formed answer and keeps moving, whether or not that answer is right.
“It didn't error” is not one failure mode; it is a description that covers at least three different mechanisms, and they call for different fixes.
The distinction has a precise name in the standards literature Canadian privacy regulators point to when assessing AI tools. The U.S. National Institute of Standards and Technology's AI Risk Management Framework defines reliability as the “ability of an item to perform as required, without failure, for a given time interval, under given conditions” — and adds that “Deployment of AI systems which are inaccurate, unreliable, or poorly generalized to data and settings beyond their training creates and increases negative AI risks and reduces trustworthiness.” Note the phrase perform as required. An automation that keeps executing without an exception has satisfied the “without failure” half of that definition. It has said nothing about the “as required” half. Only one half of reliability throws an alarm on its own; the other half has to be checked for.
Worked example
An illustrative scenario, not a reported case. A brokerage runs an AI intake step that reads incoming lead forms and files each one into the CRM under the right service line. It works cleanly for months. Then a referral partner's website is redesigned, and a field that used to read “preferred contact time” is relabelled “best time to reach you.” The AI step does not error — it still produces a category for every single lead that arrives. It simply starts misfiling a slice of them, because the phrasing it was reading a signal from has moved. Nothing in the pipeline says so, because nothing was built to check whether “still runs” and “still correct” had quietly come apart. The first anyone hears of it is a client complaint weeks later — not an alert.
Not a better model — a habit of checking the model's ongoing effectiveness against the job it was set to do, on a schedule, rather than assuming that a step which was correct at launch stays correct indefinitely. The Office of the Privacy Commissioner's generative-AI principles put a version of this requirement on the table directly: organizations should “Evaluate the validity and reliability of the generative AI tool for the intended purpose” on an ongoing basis, because “the tool should be more than simply potentially useful” and “this consideration should be evidence-based.” The OPC's own footnote on that requirement points straight at the NIST framework above — a U.S. framework, cited by a Canadian regulator, for exactly this reason.
Canada's federal voluntary code for advanced generative systems names the same idea as one of its committed outcomes, worth stating in full because it describes the fix rather than just the risk: “Human Oversight and Monitoring – System use is monitored after deployment, and updates are implemented as needed to address any risks that materialize” (ISED’s voluntary code of conduct). It is voluntary and binds only its signatories, not every business running an AI-touched workflow — but the outcome it names is the practical answer to a quiet failure: somebody has to be watching after launch, not only at launch.
Deciding what to monitor, how often, and who owns the correction when the monitoring finds something is the operational half of this problem — covered in what AI can and cannot do today and, in more depth, on the day-to-day operations side of the academy linked below.
No. It means the process has not stopped, which is a narrower claim than “the process is producing correct output.” A generative step almost always returns a well-formed answer; whether that answer is right has to be checked separately, because the step itself has no built-in way to refuse a plausible-looking wrong answer the way a script refuses a malformed input.
Yes — a mis-scoped filter or a regular expression that silently stops matching anything is a long-standing example. The difference is how often it happens and how confident the wrong output looks. A broken regex usually returns nothing, which is at least a visible symptom. A model-driven step returns something fluent, which reads as a success unless someone checks it against the actual outcome.
Usually a shift in the distribution of outputs — a category that used to appear occasionally suddenly appearing constantly, or the reverse — or a disagreement between the automation's output and a small manual audit sample. Neither shows up unless someone is deliberately sampling and comparing on a schedule, which is the practical content of being “monitored after deployment”.
The mechanics of building that monitoring layer — what to sample, how often, and who owns the fix when it finds a drift — are the day-to-day work of running an AI system after it goes live.