Treadstone Associates
Article · 8 min read

Why automations break quietly

An ordinary automation tells you when it's broken. An AI-touched one often doesn't — it keeps producing an answer, and the answer looks fine right up until someone checks it against reality.

Treadstone Associates · Updated 2026

Key takeaways

  • • Rule-based automation fails by stopping — a mismatched input throws an error and someone notices. An AI-touched step rarely stops; it produces something fluent instead, and fluent is not the same as correct.
  • • Three real mechanisms behind a quiet failure: the input's shape drifting underneath the automation, an output that is wrong without being obviously wrong, and a dataset or feed that has been degraded upstream of the step doing the work.
  • • Canada's Cyber Centre states plainly that generative output “can be incorrect” and “might not take certain factors into account” — not as an edge case, but as a standing property of the technology.
  • • “Still running” and “still correct” are two different claims, and only one of them sets off an alarm. Closing the gap is a monitoring problem, not a smarter-model problem.

Two different shapes of failure

Traditional, rule-based automation tends to fail loudly. A script expects an invoice number in a fixed format; feed it something else and it throws an exception, the job queue backs up, and somebody gets paged. The failure is inconvenient, but it is visible — the system stops doing the wrong thing because it stops doing anything at all.

A step built on a generative model rarely behaves that way. Ask it to classify a lead, summarize a document or draft a response, and it will almost always return something — fluent, on-topic, and structurally correct-looking. Canada's Canadian Centre for Cyber Security describes generative AI as a technology that “generates new content by modelling features from large datasets that were fed into the model” rather than one that checks its output against a fixed rule. There is no built-in equivalent of the exception the script throws. The pipeline downstream receives a well-formed answer and keeps moving, whether or not that answer is right.

Where the quiet actually comes from

“It didn't error” is not one failure mode; it is a description that covers at least three different mechanisms, and they call for different fixes.

  1. The input's shape drifts underneath the automation. A partner's web form renames a field, a vendor changes an export format, a CRM adds a column. A rigid script trips over the change immediately. A model-driven step often absorbs it, producing a plausible answer from the altered input without any signal that the ground has shifted.
  2. The output is wrong without being obviously wrong. The Cyber Centre's guidance on generative AI is direct about this: an output "can be incorrect," "might not make sense," "might not take certain factors into account," or "can be biased" (ITSAP.00.041). None of those states looks like a crash. Each one looks like an answer.
  3. The feed itself has been degraded upstream. The same guidance flags "poisoned datasets" as a named risk: “Threat actors can inject malicious code into the dataset used to train the generative AI system”…“could also increase the potential for large-scale supply-chain attacks.” The threat-actor framing is about training data specifically, but the underlying logic generalizes to any automation that keeps learning from, or simply keeps consuming, a live external feed: what goes in shapes what comes out, and a degraded input does not announce itself.

“Still runs” is not “still correct”

The distinction has a precise name in the standards literature Canadian privacy regulators point to when assessing AI tools. The U.S. National Institute of Standards and Technology's AI Risk Management Framework defines reliability as the “ability of an item to perform as required, without failure, for a given time interval, under given conditions” — and adds that “Deployment of AI systems which are inaccurate, unreliable, or poorly generalized to data and settings beyond their training creates and increases negative AI risks and reduces trustworthiness.” Note the phrase perform as required. An automation that keeps executing without an exception has satisfied the “without failure” half of that definition. It has said nothing about the “as required” half. Only one half of reliability throws an alarm on its own; the other half has to be checked for.

Worked example

An illustrative scenario, not a reported case. A brokerage runs an AI intake step that reads incoming lead forms and files each one into the CRM under the right service line. It works cleanly for months. Then a referral partner's website is redesigned, and a field that used to read “preferred contact time” is relabelled “best time to reach you.” The AI step does not error — it still produces a category for every single lead that arrives. It simply starts misfiling a slice of them, because the phrasing it was reading a signal from has moved. Nothing in the pipeline says so, because nothing was built to check whether “still runs” and “still correct” had quietly come apart. The first anyone hears of it is a client complaint weeks later — not an alert.

What actually closes the gap

Not a better model — a habit of checking the model's ongoing effectiveness against the job it was set to do, on a schedule, rather than assuming that a step which was correct at launch stays correct indefinitely. The Office of the Privacy Commissioner's generative-AI principles put a version of this requirement on the table directly: organizations should “Evaluate the validity and reliability of the generative AI tool for the intended purpose” on an ongoing basis, because “the tool should be more than simply potentially useful” and “this consideration should be evidence-based.” The OPC's own footnote on that requirement points straight at the NIST framework above — a U.S. framework, cited by a Canadian regulator, for exactly this reason.

Canada's federal voluntary code for advanced generative systems names the same idea as one of its committed outcomes, worth stating in full because it describes the fix rather than just the risk: “Human Oversight and Monitoring – System use is monitored after deployment, and updates are implemented as needed to address any risks that materialize” (ISED’s voluntary code of conduct). It is voluntary and binds only its signatories, not every business running an AI-touched workflow — but the outcome it names is the practical answer to a quiet failure: somebody has to be watching after launch, not only at launch.

Deciding what to monitor, how often, and who owns the correction when the monitoring finds something is the operational half of this problem — covered in what AI can and cannot do today and, in more depth, on the day-to-day operations side of the academy linked below.

Common questions

If an AI-touched process never throws an error, does that mean it's working correctly?

No. It means the process has not stopped, which is a narrower claim than “the process is producing correct output.” A generative step almost always returns a well-formed answer; whether that answer is right has to be checked separately, because the step itself has no built-in way to refuse a plausible-looking wrong answer the way a script refuses a malformed input.

Can ordinary, non-AI automation fail quietly too?

Yes — a mis-scoped filter or a regular expression that silently stops matching anything is a long-standing example. The difference is how often it happens and how confident the wrong output looks. A broken regex usually returns nothing, which is at least a visible symptom. A model-driven step returns something fluent, which reads as a success unless someone checks it against the actual outcome.

Without an error message, what is the first real sign something has gone quietly wrong?

Usually a shift in the distribution of outputs — a category that used to appear occasionally suddenly appearing constantly, or the reverse — or a disagreement between the automation's output and a small manual audit sample. Neither shows up unless someone is deliberately sampling and comparing on a schedule, which is the practical content of being “monitored after deployment”.

Where this goes next

The mechanics of building that monitoring layer — what to sample, how often, and who owns the fix when it finds a drift — are the day-to-day work of running an AI system after it goes live.