Treadstone Associates
Article · 8 min read

What AI does not fix

Most AI discussion focuses on what the technology can do. The more useful question for a business evaluating it is the opposite one: which problems does adopting an AI tool leave completely untouched, no matter how well the tool itself performs.

Treadstone Associates · Updated 2026

Key takeaways

  • • Accuracy and robustness are formally in tension, not the same property — a system tuned for one can degrade on the other, per the U.S. NIST AI Risk Management Framework.
  • • Canada’s own federal-provincial privacy principles link directly to that same framework, and require evaluating a tool’s validity for “the intended purpose,” not treating a general capability as evidence of fitness for a specific task.
  • • A framework’s own governance step — deciding “whether the AI system achieves its intended purposes” — is a human decision the framework structures but does not make for you.
  • • AI does not fix a process that was never clearly defined in the first place — a tool can only reliably replace a step someone can already describe precisely.

There is a specific, recurring failure pattern in AI adoption: a tool performs well in a demo or a pilot, gets rolled out more broadly, and then produces results nobody trusts. The gap is rarely that the tool got worse. It is usually that the tool was never solving the thing it was asked to solve in production — a distinction that two separate, verifiable frameworks make explicit.

Accuracy and robustness are not the same claim

The NIST AI Risk Management Framework’s own characteristics page states that “accuracy and robustness contribute to the validity and trustworthiness of AI systems, and can be in tension with one another.” It defines accuracy as “closeness of results of observations, computations, or estimates to the true values or the values accepted as being true,” and reliability, separately, as the “ability of an item to perform as required, without failure, for a given time interval, under given conditions.” A model tuned to score well on a benchmark can be less robust to the messy, out-of-distribution inputs a real business actually produces — and a vendor demo, by construction, rarely surfaces that gap. This is a United States government framework; it is referenced here because Canada’s own privacy regulators point directly to it, discussed next.

The Canadian anchor for this exact point

Canada’s federal, provincial and territorial privacy commissioners jointly require, as Principle 3 of their generative AI guidance, that organizations “evaluate the validity and reliability of the generative AI tool for the intended purpose” — and the footnote attached to that sentence points directly to the NIST characteristics page above. The same principle adds that “tools must be accurate throughout the intended lifecycle of the tool and across the variety of circumstances in which they are used.” The Canadian regulatory language is explicit that validity is tied to a specific, intended purpose — not a general claim that “the model is accurate,” which is not, on its own, a testable statement.

It does not tell you whether to proceed — it tells you how to ask

The NIST framework’s own governance structure names four functions — Govern, Map, Measure and Manage — and its Manage function includes a step that is easy to skip in practice: “a determination is made as to whether the AI system achieves its intended purposes and stated objectives and whether its development or deployment should proceed”. That determination is a human judgment call the framework structures but explicitly does not make on your behalf. Likewise its Measure function requires that “the risks or trustworthiness characteristics that will not – or cannot – be measured are properly documented” — an instruction to write down what you are choosing not to check, not a mechanism that checks it for you.

The Canadian principle behind the same idea

The OPC’s Principle 3 also states the underlying test plainly: use of a generative AI system should be “evidence-based and establish that the tool is both necessary and likely to be effective in achieving the specified purpose” — not simply “potentially useful.” That standard rules out a common shortcut: adopting a tool because it is capable in general, rather than because someone has evidence it works for the specific task in front of it.

A tool cannot fix a process nobody has actually described

Underneath both frameworks sits a plainer, structural limit: an AI tool automates a step someone can already state precisely. If the underlying process is genuinely undefined — different staff handle an exception differently, or the “right” outcome depends on unwritten judgment calls — there is no version of the task an AI tool can be evaluated against, because there is no stable target to measure it against in the first place. This is not a limitation of any particular model; it is a property of what evaluation requires. Cleaning up the process definition is, in that sense, a precondition for judging whether AI has done anything, not a side project.

What still requires a person

Both frameworks converge on the same residual point: oversight does not disappear once a tool is deployed. Canada’s own ISED Voluntary Code names, as a manager’s obligation, to “monitor the operation of the system for harmful uses or impacts after it is made available… including through the use of third-party feedback channels”, and the OPC principles require organizations to provide impacted individuals “an effective challenge mechanism for any administrative or otherwise significant decision made about them… and allowing them the opportunity to request human review.” A validity test run once at launch, with no ongoing monitoring and no route for a person to challenge a wrong output, does not satisfy either standard — and no model performance, however good, substitutes for that structure.

Where the challenge right stops being voluntary

Everything cited above — the NIST framework, the OPC principles, the ISED code — is guidance, not law; none of it can be enforced by a regulator issuing a fine. Québec’s privacy statute is the exception. Section 12.1 of the Act respecting the protection of personal information in the private sector requires a business that renders “a decision based exclusively on an automated processing” of personal information to give the affected person “the opportunity to submit observations” to staff able to review the outcome — a legal right, not a best practice. Failing to provide it can draw an administrative monetary penalty of up to $10,000,000 or 2% of worldwide turnover. That is the concrete version of the review mechanism the frameworks above only recommend: in Québec, at least, a person’s right to challenge the output is not optional once the decision is exclusively automated.

A short version of the whole point

“This model is accurate” is not a complete claim. This model is valid and reliable for this specific, defined task, evaluated with evidence, monitored after deployment, with a route for a person to challenge a wrong result is what both the U.S. technical framework and Canada’s own privacy principles actually require before AI should be trusted with a decision that affects someone.

Related: why most AI pilots stall, and what changes first when a business adopts AI.

Common questions

Does high accuracy mean an AI tool is trustworthy?

Not on its own. NIST’s own framework describes accuracy and robustness as properties that “can be in tension with one another” — a model can score well on one measure and degrade on the other for real-world inputs.

What does Canadian privacy guidance actually require before trusting a tool?

Evidence that the tool is “necessary and likely to be effective in achieving the specified purpose” — a specific, evidence-based test, not a general claim that the technology is capable.

Can AI fix a process that isn’t clearly defined yet?

No. Evaluating whether a tool performs a task correctly requires a stable definition of what “correct” means for that task — a process nobody has actually written down cannot be measured against, whatever tool is applied to it.

Trying to work out what a specific AI tool actually won’t fix for you?

A short call is enough to separate a genuine fit from a capability that won’t transfer to your process.