Treadstone Associates
Article · 8 min read

Why most AI pilots stall

“The pilot didn’t work” and “the pilot stalled” are usually two different stories being told with the same sentence. The second, far more common one is a project that never actually failed a test — because no test was defined precisely enough to fail.

Treadstone Associates · Updated 2026

Key takeaways

  • • A framework’s Map function requires that “context is established and understood” before deployment; a pilot without a defined intended purpose has skipped this step, which makes later evaluation impossible.
  • • Canada’s own privacy principles require that AI use be “evidence-based” and shown “necessary and likely to be effective” — a higher bar than “seemed promising in a demo,” and one many pilots never actually clear.
  • • Cost (10.6%) and cybersecurity or privacy concerns (13.4%) are the two most-cited barriers limiting AI use among Canadian businesses — concentrated in specific industries, not spread evenly.
  • • “Lack of skilled workers” as a barrier hits information and cultural industries (20.9%) and manufacturing (12.0%) hardest — a pilot in either sector is disproportionately likely to stall on this specific constraint.

A pilot that stalls almost never dies in one dramatic meeting. It quietly stops being anyone’s priority, the person who championed it moves on, and eighteen months later nobody can say definitively whether it worked. Two governance frameworks — one American, one Canadian — both point to the same root cause, and Statistics Canada’s barrier data shows where it hits hardest.

The step that gets skipped: defining the target before starting

The NIST AI Risk Management Framework’s Map function opens with a category stated in four words: “context is established and understood.” Its first subcategory spells out what that requires: “intended purposes, potentially beneficial uses, context-specific laws, norms and expectations, and prospective settings in which the AI system will be deployed are understood and documented.” A pilot launched to see what this tool can do has, by construction, skipped this step — there is no documented intended purpose to measure the pilot against later, which means there is no way to say afterward whether it succeeded or failed. It simply runs until attention moves elsewhere.

The Canadian bar is explicit, and higher than “promising”

Canada’s federal, provincial and territorial privacy commissioners require, in their joint generative AI principles, that use of the technology be “evidence-based and establish that the tool is both necessary and likely to be effective in achieving the specified purpose” — explicitly rejecting “simply potentially useful” as a sufficient standard. A pilot greenlit because a demo looked impressive, without a specified purpose or evidence tied to it, has not cleared this bar even if the underlying model performs well. The gap between “this looked promising” and this is evidence-based and necessary for a specified purpose is exactly where a large share of stalled pilots live.

The measurement step that documents its own gaps

The same framework’s Measure function requires that “the risks or trustworthiness characteristics that will not – or cannot – be measured are properly documented.” In practice, most pilots never write this document at all — nobody records what wasn’t checked, so there is no way, months later, to distinguish a pilot that quietly failed a test everyone forgot to run from one that never had a test in the first place. Both look identical from the outside: a project that used to be discussed in meetings and now isn’t.

Where the resourcing barriers concentrate

Even a well-scoped pilot runs into real, measured constraints, and StatCan’s own barrier data shows they are not evenly distributed. Cybersecurity or privacy concerns were the most-cited barrier limiting AI use nationally at 13.4%, followed by cost at 10.6%. Cybersecurity and privacy concerns were highest in information and cultural industries (30.9%) and health care and social assistance (26.4%) — the same two industries also led on regulatory concerns (21.9% and 17.6% respectively). Cost was the more common barrier in information and cultural industries (23.6%) and professional, scientific and technical services (14.7%). “Lack of skilled workers” as a barrier concentrated in information and cultural industries (20.9%) and manufacturing (12.0%). A pilot run inside any of these specific industries is disproportionately likely to hit the exact barrier its own sector reports most, rather than a generic, evenly spread obstacle.

Where the Canadian regulatory concern has actual teeth

Québec is the one place this barrier is not just a perception. Section 3.3 of the Act respecting the protection of personal information in the private sector requires any business to “conduct a privacy impact assessment for any project to acquire, develop or overhaul an information system...involving the collection, use, communication, keeping or destruction of personal information”, and to “consult the person in charge of the protection of personal information within the enterprise from the outset of the project” — the same accountability role PIPEDA requires outside Québec. A pilot that starts moving data into a new AI system before that consultation happens has skipped a legal step, not merely a best-practice one — a sharper version of the “context is established and understood” requirement described above, with an actual regulator behind it.

The decision step nobody wants to own

Both frameworks eventually require an explicit yes-or-no. NIST’s Manage function states that “a determination is made as to whether the AI system achieves its intended purposes and stated objectives and whether its development or deployment should proceed.” A stalled pilot is, structurally, one where this determination never gets made — not a “no,” which would end the project cleanly, but an absence of any decision at all, which lets the pilot linger indefinitely in a state that is neither adopted nor killed.

The three questions a pilot needs answered before it starts

(1) What specific, named task is this replacing or assisting with — not “see what it can do”? (2) What evidence, gathered during the pilot, would count as success or failure? (3) Who is responsible for making the proceed-or-stop determination on a specific date, rather than letting it lapse by default? A pilot that can answer all three before it launches is far less likely to be the one nobody can explain eighteen months later.

Why a stall is easy to mistake for patience

A stalled pilot rarely announces itself as a failure, which is part of why it persists. Nobody schedules a meeting to declare it dead; it simply stops appearing on the agenda, the champion who pushed for it takes on other priorities, and the absence of a documented decision means nobody is technically wrong to assume it is still “in progress.” That is a materially different problem from a pilot that ran its defined test and failed it — a failed test produces a clear next step, even if the next step is stopping. An undefined pilot produces no next step at all, which is a slower and more expensive failure mode precisely because it never looks like one from the inside.

Related: what AI does not fix, and what changes first when a business adopts AI.

How to structure that decision before committing budget to a pilot is covered on the strategy and roadmapping hub.

Common questions

Does a stalled pilot mean the AI tool doesn’t work?

Not necessarily — a stall usually means no evidence-based determination was ever made, not that the tool failed a defined test. Most stalled pilots never had a defined test to fail.

What is the most common barrier Canadian pilots actually run into?

Cybersecurity or privacy concerns, cited by 13.4% of businesses nationally and by 30.9% in information and cultural industries specifically (StatCan, Q2 2026).

What does Canadian privacy guidance say has to happen before a pilot is treated as a success?

The use has to be “evidence-based” and shown “necessary and likely to be effective” for its specified purpose — a standard many pilots never formally test against.

Have a pilot that’s quietly stalled?

A short call is enough to work out whether it needs a defined test, or a clean decision to stop.