A polished demonstration is evidence that a tool can work under the conditions its maker chose. It is not evidence that it will work under the conditions your business actually has — and the gap between the two is structural, not a matter of finding a more honest vendor.
Key takeaways
A demo runs the input that shows a product at its best, because that is how a short pitch has to work. That is not necessarily dishonest — it is simply a different exercise than the one the NIST AI Risk Management Framework describes for genuine accuracy testing, which calls for clearly defined and realistic test sets – that are representative of conditions of expected use. A single curated example, chosen in advance by the people selling the tool, is close to the opposite of that standard by construction.
This is true even when nobody involved is trying to mislead anyone. A salesperson demonstrating their own product naturally reaches for the example they know works well, the same way anyone showing off a skill picks their best trick rather than a random one. The bias toward the flattering case does not require bad faith; it is built into what a demonstration is for.
A demo processes one record, one document, one call, in the time it takes to watch. Robustness — the property that actually matters once a tool is running your work — is defined as the ability of a system to maintain its level of performance under a variety of circumstances. One pass through one example cannot demonstrate a variety of circumstances; it can only demonstrate the one circumstance chosen for the room.
Volume also surfaces problems that are statistically rare but operationally significant. A failure mode that shows up in one case out of five hundred will essentially never appear in a fifteen-minute demo, and will show up regularly the moment the tool is processing hundreds of cases a week — at which point it is no longer a curiosity, it is a recurring cost someone has to notice and handle.
Data is often cleaned, categories pre-defined and integrations pre-wired ahead of a demo — setup work that a real deployment has to do on your own systems and that simply is not visible in the fifteen minutes you watch. A tool that looks effortless in a demo can still require weeks of that groundwork before it produces anything usable on your own, messier data.
It is worth asking directly, in the room: “what did you have to do to this data before this demo, that we would also have to do to ours?” A vendor confident in the product will usually have a real answer — a data format, a minimum record quality, an integration step — and that answer is more useful to you than another minute of watching the polished version run.
Canada's privacy principles for generative AI set a bar for adopting a tool at all: an organization should establish that the tool should be more than simply potentially useful. This consideration should be evidence-based and establish that the tool is both necessary and likely to be effective in achieving the specified purpose. A demo shows a tool being potentially useful. It does not, on its own, establish the evidence-based case that principle calls for — and in particular it almost never shows what the tool does when it is unsure, or wrong, and whether it flags that instead of guessing with the same fluent confidence it uses for a correct answer.
A tool worth adopting should have a visible way of saying “I am not sure about this one” and handing the case to a person, rather than guessing with full confidence every time. That behaviour is almost never demonstrated voluntarily, because it looks, on the surface, like the tool failing — even though, from a deployment standpoint, a tool that flags its own uncertainty is doing exactly what a well-designed system should do.
Four questions travel well across almost any AI demo, regardless of what the product does: what happens with an input that is missing a field or spelled inconsistently; what does the tool do when it is not confident, and does it say so; what setup, cleanup or integration work happened before this session that we would also have to do; and can we bring our own example right now and watch it run live. A vendor's answers to those four questions tell you considerably more than the polish of the scripted portion of the demo, because they are the questions a curated demonstration is specifically built not to need answered.
Bring your own ordinary example — not your worst case, just a typical one — and watch what the tool does with it live, including the messy version of it. Put it in front of the person who would actually use it day to day, not only the manager evaluating the purchase. A step-by-step version of exactly that process, how to tell a real AI use case from a polished demo, walks through how to run this evaluation once you know what a demo is structurally unable to show you.
None of this requires distrusting the vendor. It requires treating a demo for what it structurally is — a curated, single-pass, best-case illustration — and supplying the volume, the messiness and the failure-path scrutiny that a fifteen-minute pitch was never built to provide.
See also how AI benchmarks mislead and common AI myths in business.
Not necessarily — it is how a short pitch has to work. The obligation it creates is on the buyer's side: to supply the scrutiny the demo format cannot, before treating it as evidence. A vendor who resists letting you bring your own example, however, is a separate and more useful signal.
Not on its own. Length is not the same as representativeness. What matters is whether the input tested was your own ordinary example, not a curated one, however long the session runs, and whether the messy version of that example was included rather than a cleaned-up copy of it.
The person who will use it day to day, not only the manager who evaluated it in the sales call — that person notices, within minutes, whether the tool fits how the work actually gets done, in a way a sales conversation rarely surfaces.
A vendor confident in their own product generally welcomes the questions above, because they are the same questions a serious buyer would eventually ask anyway, just asked earlier. A reluctance to answer them is itself useful information about how the tool will hold up outside the sales process.
Only if it runs on your own ordinary data, over enough volume to surface the rare cases, and with the actual future users involved. A short, vendor-run proof-of-concept on a handful of curated examples is still, structurally, a longer demo rather than a genuine evaluation.
Turning “it looked good in the demo” into a defensible short list of use cases worth testing is a scoping exercise. For that, see AI opportunity mapping: finding your first three use cases.