Treadstone Associates
Ask an Expert · 3 min read

How do you measure if AI is working?

By comparing a defined measure — time, error rate or cost on a specific task — against what the same task looked like before AI touched it, not by how polished or confident the output feels.

Treadstone Associates · Updated 2026

Short answer

By comparing a defined measure — time, error rate or cost on a specific task — against what the same task looked like before AI touched it, not by how polished or confident the output feels. That is a different question from whether one output can be trusted, or whether a system still holds up months later.

Start from a number you were already tracking

Statistics Canada’s Q2 2026 survey of AI-using Canadian businesses found that 44.4% made changes to their training or staffing practices as a direct result of adopting AI — a visible, countable change, not a vague impression that things felt different.

The method behind any such finding is the same one a single business can run: pick a number you were already measuring before AI touched the task — response time, error rate, hours logged, calls handled — and compare it to the same number afterward. Without that baseline, “did it work” has nothing to be checked against. Choosing which task to point AI at in the first place is a different, earlier question — see where AI actually pays for itself first.

Measurement is its own ongoing function, not a one-time check

The NIST AI Risk Management Framework — a U.S. standard, cited here because it is the one Canada’s own privacy guidance references for evaluating a tool’s validity — structures its whole approach around four functions: Govern, Map, Measure and Manage. Measurement is treated as its own distinct, ongoing function, separate from setting policy or managing risk once something goes wrong.

The practical takeaway for a single business is smaller than the framework itself: decide up front which number counts as success, and write it down, rather than deciding after the fact whether the result feels good enough.

What this question is not

This is different from asking whether one specific answer can be trusted — see how to tell if an AI output is trustworthy — and different again from whether a system that worked at launch still works months later, covered in why AI systems degrade over time. This one is about setting up the before-and-after comparison in the first place, so those two later questions have something concrete to be checked against.

Ready to set the comparison up properly?

See how to sequence an AI change so the measurement is built in from the start.