An agent with real permissions can do real damage as easily as real good, and its own judgment is not the thing standing between the two. That job belongs to guardrails — enforced limits, checked deliberately, not hoped for.
Key takeaways
It is tempting to treat a well-written prompt or a well-behaved model as sufficient protection against a bad outcome. A guardrail is a different kind of thing entirely: a limit enforced by something other than the model’s own output — a permission scope that makes an action impossible rather than merely discouraged, a code-level check that refuses to execute below a confidence threshold, a required human approval before a consequential step. See what an AI can reach once it’s connected for why the scope of what’s reachable is the first and most important guardrail of all.
The OPC’s generative AI principles, issued jointly with British Columbia, Quebec and Alberta’s privacy commissioners, name three concrete practices. First, avoiding what the document calls “no-go zones” — organizations should not “develop or put into service generative AI systems that violate ‘no-go zones’” — among the examples the guidance gives, “profiling that may lead to unfair, unethical, or discriminatory treatment, or creating outputs that threaten fundamental rights and freedoms.” Second, testing designed to find problems deliberately, through an adversarial or red-team testing process “to identify potential unintended inappropriate uses of the generative AI system.” Third, a necessity check before deployment at all: “consider whether the use of a generative AI system is necessary and proportionate…the tool should be more than simply potentially useful. This consideration should be evidence-based and establish that the tool is both necessary and likely to be effective in achieving the specified purpose.”
The same principles add a fourth practice that functions as a guardrail in its own right, rather than only a courtesy: labelling. The document asks organizations to “ensure that system outputs that could have a significant impact on an individual or group are meaningfully identified as being created by a generative AI tool.” A labelled output is one a reviewer knows to treat with the appropriate level of scrutiny; an unlabelled one that looks identical to human-authored content invites exactly the kind of unearned trust that lets a wrong or fabricated claim slip through unchecked.
Canada’s Cyber Centre’s generative AI guidance gives the concrete failure modes these practices exist to prevent: “users may unknowingly provide sensitive corporate data or personally identifiable information (PII) in their AI queries and prompts” once a system is connected to live data; a “poisoned” training dataset “could also increase the potential for large-scale supply-chain attacks” once a compromised model is deployed widely; and an AI coding assistant can see a developer “inadvertently introduce insecure and buggy code into the development pipeline” without a review step designed to catch it. None of these is prevented by the model simply being asked to behave — each needs an enforced check sitting outside the model itself.
“Shadow AI” — an AI tool or agent an employee has wired into their own work without it going through any formal review — is not usefully understood as a rogue or malicious system. It’s a scoping failure: nobody defined what it should be allowed to reach, so by construction nobody can say what it’s actually touching. The Cyber Centre’s “privacy of data” risk applies here with more force than in a formally reviewed deployment, precisely because the reach was never deliberately scoped in the first place — there is no record of what was decided, because nothing was decided. See what an AI can reach once it’s connected for what deliberate scoping actually looks like by contrast.
An employee connects a personal AI-agent browser extension to their work email and calendar to save time drafting replies and scheduling. Nobody in the organization reviewed what the extension can read, what it stores, or where it sends data — it simply started working, which is exactly what makes it convenient and exactly what makes it a guardrails gap. Contrast that with a formally reviewed integration built with an explicit, written scope: which mailbox, which calendar, read-only or read/write, and who signed off. Both connect an AI tool to the same two systems. Only one of them has a guardrail in place at all.
Extend the contrast one step further: the reviewed integration also logs every action it takes and labels any customer-facing message it drafts as AI-assisted before a person sends it, where the unreviewed browser extension does neither. The gap between the two setups isn’t a difference in what the underlying AI model can do — it’s entirely a difference in what was deliberately built around it before it was allowed to touch anything real. Canada’s Voluntary Code of Conduct adds a guardrail the OPC principles above don’t name: ongoing monitoring after deployment. Its Human Oversight and Monitoring commitment asks the organization managing a system to “monitor the operation of the system for harmful uses or impacts after it is made available…and inform the developer and/or implement usage controls as needed to mitigate harm.” ISED, Voluntary Code of Conduct See also Treadstone Law’s red flags in a vendor’s data practices.
Related: for what “human in the loop” adds on top of a well-scoped permission boundary, see what “human in the loop” really means; for the documented failures these guardrails exist to catch, see where AI agents still fail.
A guardrail is enforced by something outside the model’s own judgment — a permission scope, a code-level check, a required approval — where an instruction the model is simply asked to follow depends on the model actually complying every time, with nothing else backing it up if it doesn’t.
No — it usually just means an unreviewed tool, not a malfunctioning or malicious one. It’s a scoping failure — nobody defined its reach — rather than a case of the AI system behaving against its design.
The OPC’s principle is written for anyone who develops or deploys a generative AI system, not only large labs. The practice — deliberately trying to break your own setup before relying on it — scales down fine even where the formality of a dedicated red team does not.
Not as a general legal requirement — the labelling practice above comes from privacy regulators’ guidance, not a statute, and it applies specifically to outputs that could have a significant impact on a person. It is a recommended guardrail, not a universal legal obligation on every AI-generated piece of text.
Guardrails are only as good as the ongoing monitoring that checks they still hold.