Treadstone Associates
Ask an Expert · 4 min read

What happens when an AI agent makes a mistake?

Nothing catches it automatically unless someone built a check for it — the fallback has to be a designed review step, not the tool’s own confidence.

Treadstone Associates · Updated 2026

Short answer

Nothing, automatically. Canada’s own cyber-security guidance is blunt that generative AI output “can be incorrect,” so whatever catches a mistake has to be something a person deliberately built — a review step, a validation rule, a place for the error to surface — not the tool noticing it got something wrong.

What Canada’s cyber-security guidance actually says

The Canadian Centre for Cyber Security’s own generative-AI guidance lists the failure modes plainly, as a short set of warnings rather than a single reassuring sentence: outputs “can be incorrect”, “might not take certain factors into account”, and can be biased. The guidance is direct about what that means for the person relying on the output: “You should always be aware of and validate your sources to verify whether the content being presented is accurate”. Nothing in that guidance suggests the tool flags its own errors — validating the output is left to the person or process using it.

What the named organizational practice looks like

Canada’s Voluntary Code of Conduct for advanced generative AI names the practice at the organizational level. Developers who sign it commit to “maintain a database of reported incidents after deployment, and provide updates as needed to ensure effective mitigation measures”. For a business using someone else’s AI agent rather than building one, the equivalent is a log of the agent’s mistakes and a route to send a bad output back to a person — the same idea, scaled down to a single deployment rather than a whole product. The same code also names where a mistake is often first noticed in practice: managers of a publicly available system are asked to “monitor the operation of the system for harmful uses or impacts after it is made available, including through the use of third-party feedback channels, and inform the developer… as needed to mitigate harm” — in other words, a complaint from someone affected is itself part of the detection mechanism, not a separate failure of it.

A working fallback needs three parts: a way to notice the agent might be wrong — a confidence flag, a required field left blank, a customer complaint — a person who reviews before the mistake causes harm, more urgently the more consequential the action is, and a record of what happened so it does not repeat. See whether AI agents need supervision and where AI agents still fail.

Building a fallback for when an agent gets it wrong?

See how ongoing accuracy gets checked once an agent is live.