Treadstone Associates
Article · 8 min read

Can AI be wrong and still be useful?

A tool that is sometimes wrong can still be worth using. A tool that is sometimes wrong and unreviewed cannot be. The line between those two is what decides whether an imperfect AI output belongs in a workflow.

Treadstone Associates · Updated 2026

Key takeaways

  • • Usefulness is not the inverse of an error rate — it is a function of what happens after an output is produced: is it reviewed, is a wrong answer cheap to catch, and is the decision it feeds reversible.
  • • Canada's Cyber Centre states as a standing property of generative output, not an edge case, that it “can be incorrect,” “might not make sense” and “can be biased” — and that a user should “always be aware of and validate” what it produces.
  • • The OPC's own generative-AI principles make necessity conditional on evidence: a tool must be shown “necessary and likely to be effective” for its specific purpose, not simply plausible in general.
  • • No Canadian accuracy, error-rate or reliability benchmark exists for this programme to cite — the honest answer is structural: match the task's tolerance for error to the review step around the tool, not to a number nobody has published.

Usefulness is not the opposite of “sometimes wrong”

It is tempting to treat “is this AI tool accurate?” as the whole question. It isn't. A spell-checker is wrong sometimes and nobody stops using it, because a wrong suggestion is cheap to notice and cheap to reject. A single miscalculated number quietly carried into a signed financial statement is a different kind of wrong, even at a lower error rate, because nothing downstream is built to catch it. The question that actually decides whether an imperfect tool belongs in a workflow is not “how often is it wrong,” it is “what happens when it is.”

Canada has no published Canadian accuracy, error-rate or reliability benchmark for generative AI, and this article does not invent one. What does exist is a consistent, sourced description of the shape of the problem — which is enough to reason about usefulness structurally, without a number that does not exist.

What the Cyber Centre says, without hedging

The Canadian Centre for Cyber Security's guidance on generative AI does not treat occasional wrongness as a defect to be engineered away before use; it states it as a standing property to plan around. Its own words, quoted directly: outputs "can be incorrect," "might not make sense," "might not take certain factors into account," and "can be biased" (ITSAP.00.041). The guidance follows that list with an instruction, not a caveat: “you should always be aware of and validate your sources to verify whether the content being presented is accurate.” Read together, the guidance is not saying generative AI is unreliable and therefore unusable. It is saying the validation step is not optional — it is part of the tool, not an add-on to it.

The Canadian test that already exists for this: necessity and evidence

The Office of the Privacy Commissioner's principles for generative AI — developed jointly with the provincial and territorial privacy commissioners — set out a test for exactly this question, framed around whether a tool should be used at all: “the tool should be more than simply potentially useful. This consideration should be evidence-based and establish that the tool is both necessary and likely to be effective in achieving the specified purpose.” The principle also directs organizations to “Evaluate the validity and reliability of the generative AI tool for the intended purpose,” a requirement whose own footnote points to the NIST AI Risk Management Framework, a U.S. standard, as the reference for what “valid and reliable” means in practice.

NIST's own definition of accuracy is worth having on hand, because it names the tension directly: accuracy is the “closeness of results of observations, computations, or estimates to the true values or the values accepted as being true,” and, in the same document, “accuracy and robustness…can be in tension with one another in AI systems.” A system tuned to be right on the exact cases it was built for can be brittle on the ones it wasn't; a system tuned to handle a wide range of cases gracefully can lose precision on the narrow ones. Neither setting is simply “more accurate” in the abstract — it depends on what the task actually needs.

The variable that actually matters: what sits between the output and the decision

Put the Cyber Centre's list and the OPC's necessity test together and a working rule falls out: an AI output is safe to use in proportion to how cheaply a wrong version of it gets caught before it does damage.

  • Low stakes, easy to check: a first-draft summary of a long document, a suggested subject line, a rough categorization a human reviews before acting. A wrong output here costs a few seconds of correction.
  • High stakes, hard to check: a number folded silently into a report nobody re-derives, a decision made and acted on before any person looks at the reasoning behind it, an output that feeds directly into another automated step with no checkpoint in between. The same underlying error rate is far more dangerous here, because nothing stands between a wrong answer and its consequence.

Worked example

An illustrative comparison, not a reported finding. Two firms use the same drafting tool for client correspondence. One routes every draft to a person who reads it before it sends — a wrong fact gets caught and fixed in the normal course of review, the same way a junior associate's draft always got reviewed before this tool existed. The other wires the same tool to send automatically once a confidence threshold is cleared. Both firms are using a tool with the identical underlying error rate. Only one of them has actually changed what happens when it is wrong — and that difference, not the tool itself, is what determines whether the setup is safe.

This is also why “AI decides” is the wrong frame for describing a properly built process. A drafting, extracting, or summarizing step can be usefully imperfect precisely because a person still reviews and signs off on what it produces; the moment a consequential decision is made and acted on without that review, the standard the output has to meet changes completely, and the honest answer is often that no current tool should be trusted to clear it unsupervised.

This is the sorting question every AI adoption decision eventually has to answer, and it belongs to strategy rather than to the mechanism itself — covered in more depth in what AI can and cannot do today and why automations break quietly, which looks at the same question from the monitoring side rather than the deployment side.

Common questions

Is there a Canadian accuracy percentage for AI tools that businesses can rely on?

No. No Canadian regulator, standards body or statistical agency publishes an AI accuracy, error-rate or reliability benchmark, and this article does not manufacture one. What Canadian sources do publish — the Cyber Centre's guidance and the OPC's principles — describe the shape of the risk and the test for whether a tool is fit for a specific purpose, which is a structural answer rather than a numeric one.

Does a wrong AI output mean the tool shouldn't be used for that task at all?

Not necessarily — it depends on what happens next. If a wrong output is reviewed and corrected before it has any effect, the tool can still be useful even with real errors. If a wrong output flows straight into an action with nobody checking it, the same error rate becomes a much bigger problem, and that is the case where the task, not the tool, needs to change.

Who is actually responsible for validating an AI output, per Canadian privacy guidance?

The organization deploying the tool. The OPC's principles put the burden of evaluating necessity, validity and effectiveness on the organization using the system, not on the AI vendor and not on the end recipient of the output — the Cyber Centre's guidance says the same thing in plainer language: validate before you rely on it.

Where this goes next

Deciding which tasks can tolerate an occasionally-wrong AI output, and which cannot, is the first sorting question in any AI adoption plan — before a single tool is chosen.