Most AI discussion focuses on what the technology can do. The more useful question for a business evaluating it is the opposite one: which problems does adopting an AI tool leave completely untouched, no matter how well the tool itself performs.
Key takeaways
There is a specific, recurring failure pattern in AI adoption: a tool performs well in a demo or a pilot, gets rolled out more broadly, and then produces results nobody trusts. The gap is rarely that the tool got worse. It is usually that the tool was never solving the thing it was asked to solve in production — a distinction that two separate, verifiable frameworks make explicit.
The NIST AI Risk Management Framework’s own characteristics page states that “accuracy and robustness contribute to the validity and trustworthiness of AI systems, and can be in tension with one another.” It defines accuracy as “closeness of results of observations, computations, or estimates to the true values or the values accepted as being true,” and reliability, separately, as the “ability of an item to perform as required, without failure, for a given time interval, under given conditions.” A model tuned to score well on a benchmark can be less robust to the messy, out-of-distribution inputs a real business actually produces — and a vendor demo, by construction, rarely surfaces that gap. This is a United States government framework; it is referenced here because Canada’s own privacy regulators point directly to it, discussed next.
Canada’s federal, provincial and territorial privacy commissioners jointly require, as Principle 3 of their generative AI guidance, that organizations “evaluate the validity and reliability of the generative AI tool for the intended purpose” — and the footnote attached to that sentence points directly to the NIST characteristics page above. The same principle adds that “tools must be accurate throughout the intended lifecycle of the tool and across the variety of circumstances in which they are used.” The Canadian regulatory language is explicit that validity is tied to a specific, intended purpose — not a general claim that “the model is accurate,” which is not, on its own, a testable statement.
The NIST framework’s own governance structure names four functions — Govern, Map, Measure and Manage — and its Manage function includes a step that is easy to skip in practice: “a determination is made as to whether the AI system achieves its intended purposes and stated objectives and whether its development or deployment should proceed”. That determination is a human judgment call the framework structures but explicitly does not make on your behalf. Likewise its Measure function requires that “the risks or trustworthiness characteristics that will not – or cannot – be measured are properly documented” — an instruction to write down what you are choosing not to check, not a mechanism that checks it for you.
The OPC’s Principle 3 also states the underlying test plainly: use of a generative AI system should be “evidence-based and establish that the tool is both necessary and likely to be effective in achieving the specified purpose” — not simply “potentially useful.” That standard rules out a common shortcut: adopting a tool because it is capable in general, rather than because someone has evidence it works for the specific task in front of it.
Underneath both frameworks sits a plainer, structural limit: an AI tool automates a step someone can already state precisely. If the underlying process is genuinely undefined — different staff handle an exception differently, or the “right” outcome depends on unwritten judgment calls — there is no version of the task an AI tool can be evaluated against, because there is no stable target to measure it against in the first place. This is not a limitation of any particular model; it is a property of what evaluation requires. Cleaning up the process definition is, in that sense, a precondition for judging whether AI has done anything, not a side project.
Both frameworks converge on the same residual point: oversight does not disappear once a tool is deployed. Canada’s own ISED Voluntary Code names, as a manager’s obligation, to “monitor the operation of the system for harmful uses or impacts after it is made available… including through the use of third-party feedback channels”, and the OPC principles require organizations to provide impacted individuals “an effective challenge mechanism for any administrative or otherwise significant decision made about them… and allowing them the opportunity to request human review.” A validity test run once at launch, with no ongoing monitoring and no route for a person to challenge a wrong output, does not satisfy either standard — and no model performance, however good, substitutes for that structure.
Everything cited above — the NIST framework, the OPC principles, the ISED code — is guidance, not law; none of it can be enforced by a regulator issuing a fine. Québec’s privacy statute is the exception. Section 12.1 of the Act respecting the protection of personal information in the private sector requires a business that renders “a decision based exclusively on an automated processing” of personal information to give the affected person “the opportunity to submit observations” to staff able to review the outcome — a legal right, not a best practice. Failing to provide it can draw an administrative monetary penalty of up to $10,000,000 or 2% of worldwide turnover. That is the concrete version of the review mechanism the frameworks above only recommend: in Québec, at least, a person’s right to challenge the output is not optional once the decision is exclusively automated.
A short version of the whole point
“This model is accurate” is not a complete claim. This model is valid and reliable for this specific, defined task, evaluated with evidence, monitored after deployment, with a route for a person to challenge a wrong result is what both the U.S. technical framework and Canada’s own privacy principles actually require before AI should be trusted with a decision that affects someone.
Related: why most AI pilots stall, and what changes first when a business adopts AI.
Not on its own. NIST’s own framework describes accuracy and robustness as properties that “can be in tension with one another” — a model can score well on one measure and degrade on the other for real-world inputs.
Evidence that the tool is “necessary and likely to be effective in achieving the specified purpose” — a specific, evidence-based test, not a general claim that the technology is capable.
No. Evaluating whether a tool performs a task correctly requires a stable definition of what “correct” means for that task — a process nobody has actually written down cannot be measured against, whatever tool is applied to it.
A short call is enough to separate a genuine fit from a capability that won’t transfer to your process.