Treadstone Associates
Article · 9 min read

How AI benchmarks mislead

A benchmark score tells you how a model performed on someone else’s test, built for someone else’s purpose. It does not tell you how it will perform on your documents, your customers, or the version of the task your business actually has.

Treadstone Associates · Updated 2026

Key takeaways

  • • A benchmark is only informative if the test set resembles the conditions you will actually use the tool in — and most published ones do not say enough for you to check that.
  • • Under Canada's own voluntary AI code, benchmarking against “recognized standards” is something a developer does to its own system, not an independent check.
  • • Accuracy and robustness can trade off against each other; a model tuned to top a fixed test can be brittle everywhere else.
  • • An eval built against your own criteria is a different thing than a published benchmark, and it is the more useful of the two.

What a benchmark actually measures

The NIST AI Risk Management Framework’s own guidance on accuracy sets a standard most published benchmark claims do not meet: measures of accuracy should consider computational-centric measures (e.g., false positive and false negative rates), human-AI teaming, and demonstrate external validity (generalizable beyond the training conditions). Accuracy measurements should always be paired with clearly defined and realistic test sets – that are representative of conditions of expected use – and details about test methodology. A single leaderboard percentage, quoted without the test set or the methodology behind it, fails that standard by default — there is no way to know whether the conditions tested resemble the conditions you would deploy into.

Notice how much is packed into that one requirement: the test set has to be realistic, it has to represent the conditions of actual use, and the methodology has to be disclosed. A press-release number that reads along the lines of outperforms competitors by X points on a named benchmark typically discloses none of those three things, which means it is not simply an incomplete version of a good accuracy claim — it is missing the parts that would let a buyer judge whether it applies to them at all.

Benchmarking is something a vendor does to itself

Canada’s ISED Voluntary Code of Conduct on advanced generative AI lists, as a measure developers are asked to adopt for all advanced generative systems, an instruction to perform benchmarking to measure the model's performance against recognized standards. That framing is telling: it is a developer’s own obligation, tested against standards the developer selects, and the code does not require the results to be independently reproduced or disclosed to a buyer. A benchmark score you see in a vendor’s materials is, structurally, the result of that self-testing, not a third-party audit.

The same code does list a small number of measures that apply only to systems available for public use and only to developers — including third-party audits before release — but even those are narrower than an independent, ongoing check: they are a one-time gate before launch, not a continuing verification that the benchmark result still holds as the model, or your use of it, changes over time.

Robustness is not the same thing as accuracy

The NIST framework is explicit that these two properties can pull against each other: accuracy and robustness contribute to the validity and trustworthiness of AI systems, and can be in tension with one another. Robustness is defined separately, as the ability of a system to maintain its level of performance under a variety of circumstances. A model can be tuned until it tops a fixed benchmark and still be brittle the moment a real input drifts even slightly from what that benchmark tested — a genuinely common outcome, and one a single accuracy number cannot reveal.

A benchmark leaderboard, by its nature, measures accuracy on one fixed test rather than performance across a variety of circumstances, so it can reward exactly the kind of narrow tuning that trades away the robustness a real deployment needs. Two models can score within a point of each other on a published benchmark and behave very differently the moment either one meets a document format, a phrasing, or a language register the benchmark never included.

An eval built for your task is a different thing than a published benchmark

OpenAI’s own developer documentation describes what it calls an “eval,” a tool it gives customers for exactly this gap: an eval will test model outputs to ensure they meet style and content criteria that you specify. That is a materially different exercise from a marketing benchmark — it means writing your own test data and your own pass/fail criteria, then running a candidate model against them, rather than trusting a number the vendor produced against its own chosen standard.

What to ask for instead of a leaderboard number

The NIST AI Risk Management Framework's own categorization step requires that “AI capabilities, targeted usage, goals, and expected benefits and costs compared with appropriate benchmarks are understood” before a system is deployed — understood, not simply cited. A useful question set follows directly: what was the test set, does it look like our data, were failures disclosed as well as successes, and the risks or trustworthiness characteristics that will not – or cannot – be measured are properly documented. A vendor that cannot answer those questions is asking you to take the score on faith.

Canada’s own voluntary code for advanced generative AI names two further measures worth putting to a vendor directly, both broader than a single benchmark number: developers commit to use a wide variety of testing methods across a spectrum of tasks and contexts prior to deployment to measure performance and ensure robustness, and separately to employ adversarial testing (i.e., red-teaming) to identify vulnerabilities. A vendor who can describe its red-teaming results and the range of contexts it tested across is answering a materially more useful question than the one a single leaderboard score answers.

A worked example: two vendors, two numbers, one honest comparison

Suppose Vendor A advertises “94% accuracy on document classification” and Vendor B advertises “89% accuracy.” Read at face value, A looks better. Ask the three questions above and the picture can flip entirely: if A’s 94% was measured on clean, single-column invoices in English, and your documents are handwritten, multi-language intake forms, A’s number describes a task you do not have. If B discloses that its 89% was measured on a messy, mixed-format sample closer to your own documents, and further breaks that number down by document type rather than reporting one blended average, B’s lower headline figure is the more informative one — not because it is higher, but because it is checkable against your actual task. The number that matters is never the one printed largest in the vendor’s deck; it is the one you can trace back to a test set that looks like your problem.

See also common AI myths in business and what an AI demo hides.

Common questions

If a benchmark score looks impressive, is that still useful information?

It is one data point, worth having, but only if you can see what the test set contained. An impressive score on a test that does not resemble your task tells you the model is good at that test, not that it is good at your task.

Should a business run its own benchmark before buying a tool?

A small, representative sample of your own documents or queries, reviewed by a person who knows the right answer, is worth more than a published leaderboard number. See what an AI demo hides for how a live evaluation differs from a demo in the same way.

Does a higher aggregate score always mean fewer real-world errors?

Not necessarily. NIST's own guidance notes that accuracy measurements may include disaggregation of results for different data segments — a model can score well on average while failing consistently on a segment of cases (a particular document type, an accent, a minority category) that an aggregate score hides.

Where this goes next

Turning a benchmark question into an actual go/no-go decision is a scoping exercise. For a structured way to run that, see the three questions that predict whether an AI pilot will scale.