Two businesses can use the identical AI product and get very different results — one gets a genuinely useful tool, the other gets an expensive disappointment. The model is rarely the variable that explains the difference. What surrounds it — the task it was pointed at, the information it was given, and whether anyone checks its output — usually is.
Key takeaways
AI systems are strongest on tasks where a person can quickly recognize a good answer even if producing one from scratch would take longer — drafting a first version of a document, summarizing a long report, generating options to choose between. They are weakest on tasks where nobody downstream can easily tell whether the answer is right, because that is exactly the condition under which a fluent, wrong answer survives unnoticed. Before pointing a model at a task, the useful question is not whether AI can do this, it is whether anyone will notice if it gets this wrong before it matters.
A model has no independent access to your business's files, your customer records or last week's meeting unless someone puts that information directly into the context window, or a product has been deliberately built to fetch it first (the retrieval step covered in the previous article in this series). A common failure mode is asking a general-purpose chat tool a question that depends on internal company information it was never given, then treating a plausible-sounding but generic answer as if it were informed by specifics it never had access to. The fix is not a better model — it is supplying the right information in the request.
Canada's federal, provincial and territorial privacy commissioners, in their joint principles for generative AI, state the necessity test plainly: organizations should “consider whether the use of a generative AI system is necessary and proportionate… the tool should be more than simply potentially useful. This consideration should be evidence-based and establish that the tool is both necessary and likely to be effective in achieving the specified purpose.” (OPC, generative AI principles) That guidance was written with personal-information handling in mind, but the underlying discipline generalizes: before reaching for AI on a task, it is worth asking whether it is genuinely the right tool, not simply an available one.
The Canadian Centre for Cyber Security's own generative-AI guidance is direct about this: outputs “can be incorrect,” “might not make sense,” “might not take certain factors into account” and “can be biased,” and “you should always be aware of and validate your sources to verify whether the content being presented is accurate.” (Canadian Centre for Cyber Security, ITSAP.00.041) A verification step — a person reviewing the output against something they know to be true before it goes anywhere — is not a sign the tool failed. It is a structural requirement for using it responsibly, in exactly the same way a junior employee's first drafts are checked before they go to a client. Where the output feeds a decision about an actual person, Canada’s federal privacy law adds a specific duty on top of ordinary review: personal information used that way must be sufficiently accurate, complete, and up-to-date “to minimize the possibility that inappropriate information may be used to make a decision about the individual.” PIPEDA, Schedule 1, Principle 4.6
A model's competence is not evenly spread across every subject. It reflects whatever mix of data it was trained on, and a domain that was thin or absent in that mix produces weaker, less reliable output, even though the model will rarely say so — it will still generate a fluent-sounding answer either way. Canada's Voluntary Code of Conduct puts the responsibility for this squarely on the organizations that build these systems, committing developers to “assess and curate datasets used for training to manage data quality and potential biases.” (ISED, Voluntary Code of Conduct) As a user rather than a builder of these systems, you cannot inspect that curation directly, but you can watch for its symptoms: a model that handles a mainstream, widely-written-about topic fluently and accurately, but grows noticeably vaguer, more generic, or more confidently wrong on a narrow, specialized or Canadian-specific subject, is very often showing you where its training data thinned out. Treat unusually specific, checkable claims on a narrow topic with more scrutiny than fluent generalities on a well-covered one — the fluency of the sentence gives you no information about which category you are in.
A firm uses a generative tool to draft first-pass responses to routine customer questions, with a staff member reviewing every draft before it is sent. Six months later, the same firm expands the tool to draft responses that are auto-sent without review, to save the staff member's time. The model has not changed. What changed is condition four — the checking step was removed — and it is precisely the removal of that step, not a defect in the model, that determines whether the second use case works well or produces a visible, embarrassing mistake the first time the model is confidently wrong about something a customer actually knew the answer to.
Related: where AI fits in everyday work, why AI gets simple things wrong, and, for turning these conditions into an actual adoption plan, the AI strategy & roadmapping hub.
Rarely on its own. A more capable model can raise the ceiling on how good an answer can be, but if the task has no checkable answer, the model was never given the needed information, or nobody reviews the output, upgrading the model does not address any of those three problems.
The stakes of the task should set the weight of the check, not remove it entirely. A low-stakes draft might get a quick skim; a customer-facing or factual claim warrants closer review. The Cyber Centre's guidance does not carve out an exception for low-stakes content — it states the validation practice as a general one.
Ask whether a person with reasonable knowledge of the subject could look at the output and tell, within a few minutes, whether it is right, wrong, or needs a small fix. If the honest answer is that you are not sure and would have to redo the whole thing yourself to check, the task is a poor fit for unreviewed AI output.
Not reliably. Fluent, confident phrasing is available to the model regardless of how thin the underlying training coverage was on that specific topic, which is exactly why it is a poor signal to rely on. The safer habit is to scrutinize narrow or highly specific claims a bit harder than broad, well-known ones, rather than trusting tone.
This is one page in a plain-English series on how AI actually works and where it fits in a Canadian business.