Treadstone Associates
Ask an Expert · 4 min read

Can an AI system be hacked?

It's easy to picture an AI tool as a target only in the abstract. Canada's own cyber security agency names specific, concrete ways it happens.

Treadstone Associates · Updated 2026

Short answer

Yes. Beyond the more familiar risk of AI being used as a weapon (better phishing emails, faster malware), Canada's Cyber Centre names three attacks that target an AI system itself: poisoning what it learned from, fooling it after it's trained, and querying it to extract what it was trained on.

Three named attacks on the system itself

The Canadian Centre for Cyber Security’s guidance on AI describes three specific attack types, distinct from the generative-AI risk list covered in AI and cyber security basics. A data poisoning attack happens during training: “When poisoned (inaccurate) data is injected into the training dataset, the learning system may be taught to make mistakes” (Cyber Centre, Artificial Intelligence, ITSAP.00.040). An adversarial example happens after training: “The tool is fooled into classifying inputs incorrectly” — the Cyber Centre’s own example is a subtly altered stop sign that a self-driving vehicle misreads as a speed-limit sign.

The third is model inversion and membership inference: a threat actor queries an organisation’s data model directly, and “a model inversion attack will reveal the underlying dataset, allowing the threat actor to reproduce the training data,” while a membership inference attack “confirms if a specific data file is part of the training data” — both of which “could compromise the confidentiality and privacy of your training data and expose sensitive information” (Cyber Centre, ITSAP.00.040). Unlike the other two, this attack doesn’t need to touch the training pipeline at all — it works purely by asking the finished model the right questions.

Weapon versus target — a useful distinction

It helps to separate two different questions the same word “hacked” gets used for. Attacks that use AI as a tool — sharper phishing, deepfake impersonation, AI-assisted malware — are a threat-actor capability problem covered elsewhere in this hub. Attacks on an AI system — the three above, plus the prompt-injection technique covered at how hidden text hijacks an AI — are a different exposure, and the mitigations differ: prompt injection is defended by input sanitisation and output filtering, while data poisoning is defended by controlling and auditing what goes into training in the first place.

What this means for a business that didn't build the model

Most businesses use an AI tool someone else built and trained, which removes direct exposure to poisoning the training set — that risk sits with the vendor. It does not remove exposure to model inversion and membership inference, which target the finished, deployed system through ordinary-looking queries, or to adversarial examples crafted against whatever the tool is actually used for day to day.

Where this goes next

Keeping a deployed AI tool resistant to these attacks over time — not just at launch — is exactly the kind of standing responsibility an operations function owns.