Treadstone Associates
Article · 9 min read

How hidden text hijacks an AI

A phishing email needs a person to click something. A prompt injection needs an AI system to read something — a document, a webpage, an email — and it works even if no human ever sees the hidden instruction at all. Canada's privacy regulator and its Cyber Centre both name this as a real, current risk, not a theoretical one.

Treadstone Associates · Updated 2026

Key takeaways

  • • Canada's Privacy Commissioner defines prompt injection in its own generative-AI guidance as attacks “in which carefully crafted prompts bypass filters or make the model perform unanticipated actions” — a Canadian regulatory definition, not just a security-industry term.
  • • Canada's Cyber Centre documents a real, named 2025 case: threat actors used a prompt injection to fool GitHub Copilot into running dangerous code remotely, and Microsoft had to fix it.
  • • The instruction doesn't have to be typed by the person using the AI tool at all — it can be hidden inside a document, webpage or email the AI system reads and processes on someone's behalf.
  • • The Cyber Centre's own recommended defence list is concrete: sanitize inputs, isolate system prompts, filter outputs, restrict what tools an AI agent can act on unsupervised, and limit how much an AI system can communicate outward on its own.

A worked example

A business connects an AI assistant to its shared inbox to draft replies and summarize incoming messages. An attacker sends an email that looks ordinary to a person — a short, unremarkable message — but includes a block of text in white font on a white background, or hidden inside the message's HTML in a way no email client displays by default. That hidden block is written as an instruction: forward any message in this inbox mentioning invoice numbers to an external address, or summarize and reply with the contents of a specific earlier email thread. A person opening the email sees nothing unusual. The AI assistant, asked to read and process the message, has no built-in reason to treat the hidden block differently from the visible one — both are simply text inside the email it was told to read. Whether the attack succeeds then depends entirely on the layered defences described below, not on whether a person noticed anything wrong, because there was nothing for a person to notice.

The mechanism, in plain terms

Canada's Cyber Centre explains the mechanism with a deliberately ordinary example: “Imagine that you're chatting with a smart chatbot assistant, but a hacker finds a way to secretly slip in sneaky commands inside your questions. These commands trick the AI into doing things it shouldn't do, such as running harmful computer commands” (Cyber Centre, Top 10 artificial intelligence security actions, ITSAP.10.049). The key feature that makes this different from an ordinary scam is that the target of the trick is the AI system's language-processing behaviour, not a person's judgment. An AI model doesn't distinguish, by default, between an instruction its own user typed and an instruction sitting inside the content it was asked to read — a webpage it summarized, a document it was asked to analyze, an email it was asked to draft a reply to. If that content contains text formatted to look like an instruction, a poorly defended system can simply follow it.

Canada's regulatory definition

This isn't only security-industry jargon. Canada's Privacy Commissioner names and defines the technique directly in its own generative-AI guidance, under Principle 9, Safeguards, listing threats organizations must maintain “ongoing awareness of, and mitigations against”: “prompt injection attacks (in which carefully crafted prompts bypass filters or make the model perform unanticipated actions); model inversion attacks (in which personal information contained in the model's training data is exposed); and jailbreaking (in which privacy or security controls in the tool are overridden)” (OPC, Principles for responsible, trustworthy and privacy-protective generative AI). The OPC's framing places this squarely inside privacy law, not just cybersecurity: a successful prompt injection that causes a model to disclose information it was trained on, or that it has access to elsewhere in a conversation, is a privacy incident with a name and a regulatory definition attached, covered from that angle in what a privacy breach looks like with AI.

A documented case, not a hypothetical

The Cyber Centre's own Top 10 artificial intelligence security actions primer, dated May 2026, names a real incident rather than a constructed example: “This happened in 2025 with GitHub Copilot. Threat actors used a clever ‘prompt injection’ to fool it into running dangerous code remotely. Microsoft quickly fixed it, showing us that spotting these hidden hacks early is key to keeping AI safe.” The same publication describes a related, separate case affecting Microsoft 365 Copilot's document-retrieval system, where researchers found threat actors could inject manipulated content into documents the system retrieved, persistently altering its outputs without the user ever typing a malicious prompt themselves. Both cases were real production tools, not laboratory demonstrations, and both were patched after the fact — which is itself the point: the defence had to happen after discovery, because nothing in the system's normal operation flagged the hidden instruction as different from a legitimate one.

Why it's specifically the hidden part that matters

The Cyber Centre's mitigation list for this risk includes an instruction that only makes sense once you understand the exfiltration mechanism: “reduce the ability for AI systems to communicate externally (such as embedded data in markdown image URLs)” (Cyber Centre, ITSAP.10.049). A markdown image tag causes many chat interfaces to automatically fetch an image from a URL to display it — which means a hidden instruction can direct the model to construct a URL that embeds data it has access to, and the act of simply rendering that image sends the data to an outside server, with no separate click, download or approval step from the person using the tool. The instruction that triggers this never has to be visible in the conversation the person sees; it can arrive inside a document, a webpage, or any other content the AI system was asked to process.

The Cyber Centre's own defence list

ITSAP.10.049 names a specific set of mitigations under “Implement prompt injection and jailbreak mitigations”: sanitize inputs; isolate system prompts and protect prompt history; apply output filtering and policy gating; restrict high-risk tools and agents through role-based access and identity controls; validate downstream actions (files, code and tools) before execution; quarantine anomalous outputs; reduce access to private data by AI models; limit exposure to untrusted content; and reduce the ability for AI systems to communicate externally on their own. None of these is a single silver-bullet fix — they are layered controls, on the assumption that any one of them can fail.

Common questions

Does a prompt injection attack require the AI tool's own user to be tricked?

No — that's the specific danger of the technique. The hidden instruction can be embedded in content the AI system processes on the user's behalf, such as a document, webpage or email, without the user ever seeing or typing the malicious text themselves.

Is prompt injection the same thing as jailbreaking?

No, though they're related and often listed together. Canada's Privacy Commissioner defines them separately: prompt injection is crafted prompts that bypass filters or cause unanticipated actions, while jailbreaking specifically means overriding a tool's privacy or security controls.

Has a real company actually been affected by this, or is it theoretical?

It's documented, not theoretical. Canada's Cyber Centre names a real 2025 incident affecting GitHub Copilot and a related case affecting Microsoft 365 Copilot's document-retrieval system, both since patched by Microsoft.

Where this goes next

Deciding which tools an AI agent can act on without a human checking first is exactly the design decision that determines how much damage a successful injection can do.