A benchmark, in AI, is a fixed, repeatable test — the same set of tasks, questions or files, checked against the same known correct answers — used to compare how different AI systems, or different versions of the same system, perform under identical conditions.
The United States’ NIST AI Risk Management Framework treats benchmarking as one part of a wider discipline it calls measurement. “The MEASURE function employs quantitative, qualitative, or mixed-method tools, techniques, and methodologies to analyze, assess, benchmark, and monitor AI risk and related impacts”, and the framework expects that work to include “rigorous software testing and performance assessment methodologies with associated measures of uncertainty, comparisons to performance benchmarks” — not a single number quoted once and left unchecked.
That expectation has a Canadian anchor. Canada’s federal, provincial and territorial privacy commissioners direct developers to “evaluate the validity and reliability of the generative AI tool for the intended purpose”, on the basis that “tools must be accurate throughout the intended lifecycle of the tool and across the variety of circumstances in which they are used”. A benchmark that a vendor runs once, on its own favourable case, does not meet that bar on its own.
Before switching underwriting-support tools, a firm might set aside 50 real, already-closed files it never shows the vendor in advance, run both the old and new tool against that same set, and have a person check how many each one handled correctly without needing a fix. Whichever tool the firm never tested this way is the one it is trusting on faith, not on a benchmark. Repeating the same held-out set every time a tool is updated is what turns a one-off comparison into an ongoing check.
See also: what is a proof of concept, what is an AI model.
Deciding what to test before committing further, and against what standard, sits ahead of any build — ai-strategy-roadmapping covers how a firm sets that bar before choosing a direction.