A phishing email needs a person to click something. A prompt injection needs an AI system to read something — a document, a webpage, an email — and it works even if no human ever sees the hidden instruction at all. Canada's privacy regulator and its Cyber Centre both name this as a real, current risk, not a theoretical one.
Key takeaways
A business connects an AI assistant to its shared inbox to draft replies and summarize incoming messages. An attacker sends an email that looks ordinary to a person — a short, unremarkable message — but includes a block of text in white font on a white background, or hidden inside the message's HTML in a way no email client displays by default. That hidden block is written as an instruction: forward any message in this inbox mentioning invoice numbers to an external address, or summarize and reply with the contents of a specific earlier email thread. A person opening the email sees nothing unusual. The AI assistant, asked to read and process the message, has no built-in reason to treat the hidden block differently from the visible one — both are simply text inside the email it was told to read. Whether the attack succeeds then depends entirely on the layered defences described below, not on whether a person noticed anything wrong, because there was nothing for a person to notice.
Canada's Cyber Centre explains the mechanism with a deliberately ordinary example: “Imagine that you're chatting with a smart chatbot assistant, but a hacker finds a way to secretly slip in sneaky commands inside your questions. These commands trick the AI into doing things it shouldn't do, such as running harmful computer commands” (Cyber Centre, Top 10 artificial intelligence security actions, ITSAP.10.049). The key feature that makes this different from an ordinary scam is that the target of the trick is the AI system's language-processing behaviour, not a person's judgment. An AI model doesn't distinguish, by default, between an instruction its own user typed and an instruction sitting inside the content it was asked to read — a webpage it summarized, a document it was asked to analyze, an email it was asked to draft a reply to. If that content contains text formatted to look like an instruction, a poorly defended system can simply follow it.
This isn't only security-industry jargon. Canada's Privacy Commissioner names and defines the technique directly in its own generative-AI guidance, under Principle 9, Safeguards, listing threats organizations must maintain “ongoing awareness of, and mitigations against”: “prompt injection attacks (in which carefully crafted prompts bypass filters or make the model perform unanticipated actions); model inversion attacks (in which personal information contained in the model's training data is exposed); and jailbreaking (in which privacy or security controls in the tool are overridden)” (OPC, Principles for responsible, trustworthy and privacy-protective generative AI). The OPC's framing places this squarely inside privacy law, not just cybersecurity: a successful prompt injection that causes a model to disclose information it was trained on, or that it has access to elsewhere in a conversation, is a privacy incident with a name and a regulatory definition attached, covered from that angle in what a privacy breach looks like with AI.
The Cyber Centre's own Top 10 artificial intelligence security actions primer, dated May 2026, names a real incident rather than a constructed example: “This happened in 2025 with GitHub Copilot. Threat actors used a clever ‘prompt injection’ to fool it into running dangerous code remotely. Microsoft quickly fixed it, showing us that spotting these hidden hacks early is key to keeping AI safe.” The same publication describes a related, separate case affecting Microsoft 365 Copilot's document-retrieval system, where researchers found threat actors could inject manipulated content into documents the system retrieved, persistently altering its outputs without the user ever typing a malicious prompt themselves. Both cases were real production tools, not laboratory demonstrations, and both were patched after the fact — which is itself the point: the defence had to happen after discovery, because nothing in the system's normal operation flagged the hidden instruction as different from a legitimate one.
The Cyber Centre's mitigation list for this risk includes an instruction that only makes sense once you understand the exfiltration mechanism: “reduce the ability for AI systems to communicate externally (such as embedded data in markdown image URLs)” (Cyber Centre, ITSAP.10.049). A markdown image tag causes many chat interfaces to automatically fetch an image from a URL to display it — which means a hidden instruction can direct the model to construct a URL that embeds data it has access to, and the act of simply rendering that image sends the data to an outside server, with no separate click, download or approval step from the person using the tool. The instruction that triggers this never has to be visible in the conversation the person sees; it can arrive inside a document, a webpage, or any other content the AI system was asked to process.
ITSAP.10.049 names a specific set of mitigations under “Implement prompt injection and jailbreak mitigations”: sanitize inputs; isolate system prompts and protect prompt history; apply output filtering and policy gating; restrict high-risk tools and agents through role-based access and identity controls; validate downstream actions (files, code and tools) before execution; quarantine anomalous outputs; reduce access to private data by AI models; limit exposure to untrusted content; and reduce the ability for AI systems to communicate externally on their own. None of these is a single silver-bullet fix — they are layered controls, on the assumption that any one of them can fail.
No — that's the specific danger of the technique. The hidden instruction can be embedded in content the AI system processes on the user's behalf, such as a document, webpage or email, without the user ever seeing or typing the malicious text themselves.
No, though they're related and often listed together. Canada's Privacy Commissioner defines them separately: prompt injection is crafted prompts that bypass filters or cause unanticipated actions, while jailbreaking specifically means overriding a tool's privacy or security controls.
It's documented, not theoretical. Canada's Cyber Centre names a real 2025 incident affecting GitHub Copilot and a related case affecting Microsoft 365 Copilot's document-retrieval system, both since patched by Microsoft.
Deciding which tools an AI agent can act on without a human checking first is exactly the design decision that determines how much damage a successful injection can do.