Ninety days is long enough to get a real answer and short enough that a bad guess doesn't become a bad habit. Here is how to scope one so the answer is trustworthy.
Key takeaways
STEP 01 OF 10
Most of what fails in a first pilot fails because the scope was "AI in the office" rather than one bounded task. StatCan found that among Canadian businesses already using AI, adoption clusters around a short list of applications — data analytics (36.6%), text analytics (34.5%), virtual agents or chat bots (28.2%) and natural language processing (27.0%) led the list in the second quarter of 2026. Adoption succeeds task by task, not company-wide. Pick the one task in your firm with the clearest before and after: a specific step in estimating, a specific document-intake queue, a specific scheduling handoff.
Write the task down as a single sentence with a single owner and a single measurable output. If the sentence needs "and" in the middle, split it into two pilots and run the second one later.
STEP 02 OF 10
A pilot without a named owner drifts until someone senior notices it either quietly failed or quietly became load-bearing without a decision ever being made. The Voluntary Code of Conduct that Canada's innovation ministry published for advanced generative AI names accountability as its first outcome: an organisation "understands its role" and "puts in place appropriate risk management systems." Translate that into one person's name on the pilot, with the explicit authority to stop it on day 40 if the early numbers say so.
That person also owns the vendor relationship for the ninety days — one point of contact, so a support ticket or a contract question does not bounce between whoever happened to answer the phone.
STEP 03 OF 10
Before the tool is switched on, write down what it will see: pricing data, client or tenant personal information, subcontractor banking details, unpublished bid numbers. This is a full exercise on its own — see data classification before you adopt AI — but at minimum, name the categories in the pilot's one-page brief before day one, not after a question comes up in week six.
In practice this means listing the specific fields or documents the tool will read, next to who currently has access to them. If the pilot would give the tool access to something fewer than three people in the firm currently see, that is worth a second look before day one, not a note for later.
STEP 04 OF 10
Do not start the pilot's clock until you have written down what "before" looked like, using records your firm already keeps — not a number a vendor suggests. No published Canadian figure exists for what an AI rollout returns in construction; the method for building your own is covered in full in measuring return on an AI rollout. A pilot that skips this step cannot answer its own question on day 90, no matter how it felt to use.
A before-number does not need to be elaborate. A week of manual timing on the specific task, written down by the person who actually does it, beats a polished dashboard built after the fact from memory. Capture it before anyone in the pilot has seen the new tool in action, or the baseline will already be contaminated by expectation.
STEP 05 OF 10
Canada's federal, provincial and territorial privacy commissioners frame this as necessity and proportionality: evaluating whether a tool is "necessary and likely to be effective" is meant to be evidence-based, and that evidence only exists if a person is checking the output against reality the whole time, not just at the start when everyone is paying attention. Put a name against every output for the full ninety days, not the first two weeks.
The sign-off should take a specific, visible form — an initial on a document, a checkbox in the workflow, a line in a log — not an informal understanding that "someone is keeping an eye on it." An informal understanding is the first thing that quietly disappears once the tool starts feeling routine, usually around week five or six.
STEP 06 OF 10
The tool runs alongside the existing manual step for the first month without replacing anything — you are checking whether its output agrees with what a person would have done. Only once that agreement is consistent does it take over the step, and even then the sign-off from the previous step stays in place for the remaining eight weeks. A pilot that removes human review the moment it goes live is not measuring the tool, it is removing the only control that would tell you if it is wrong.
Log every disagreement between the tool and the person during the shadow month, not just the count of them. The pattern in what it gets wrong — a specific document type, a specific clause, a specific supplier's format — tells you more about whether it is ready than an overall agreement percentage does on its own.
STEP 07 OF 10
StatCan asked Canadian businesses what actually limits their AI use: cybersecurity or privacy concerns (13.4% nationally, rising to 30.9% in information and cultural industries), and cost (10.6% nationally). Neither is construction-specific, but both are worth naming out loud in the pilot brief so that if either one shows up in week five, it is a known risk being tracked, not a surprise that stalls the whole exercise.
Assign each named risk to a person who would actually notice it early — the office manager for a cost overrun, whoever handles IT for a security concern — rather than leaving the whole list owned by the pilot lead alone. A risk nobody is specifically watching for tends to surface only once it has already become a problem.
STEP 08 OF 10
Kill it, extend it another quarter, or roll it into a full adoption plan. Write the decision down with the numbers that produced it, even if the answer is "not yet." A pilot that ends without a written decision has a way of quietly restarting from zero a year later, at the same cost, having taught the firm nothing the second time either.
Hold the day-90 review as an actual meeting with the before-numbers on the table, not an email asking whether people liked the tool. Liking a tool and it having moved the number you set out to move in step four are two different questions, and only one of them belongs in the decision.
STEP 09 OF 10
Before day one, put a number on what the pilot costs if it fails outright: training hours at a loaded rate, the licence for the pilot window, and the estimator time spent shadow-testing. That figure is what the day-90 decision in step eight is actually measured against, not a vague sense of "was it worth it."
A pilot with no stated downside tends to get judged on how it felt to use, which is exactly the failure mode step eight is meant to prevent. Write the downside number next to the before-number from step four, before either one has a chance to be revised in hindsight.
STEP 10 OF 10
A pilot that only configures an existing tool is different from one that needs custom integration — connecting a takeoff tool to your accounting system, for instance. For that second kind, the National Research Council's IRAP AI Assist program exists specifically to help Canadian SMEs "develop and adapt generative AI (GenAI) and deep learning (DL) solutions," backed by $100 million over five years announced in Budget 2024, with an industrial technology advisor as the first point of contact.
Most estimating and document-intake pilots at a firm this size will not need that path — a configured off-the-shelf tool is enough. Know the option exists before assuming the whole build cost sits on your own budget, especially if the pilot's scope from step one turns out to need custom work partway through.
Scoping the pilot to a whole department instead of one task. A pilot titled "AI in estimating" cannot fail cleanly or succeed clearly, because nobody agreed in advance what it was supposed to prove. Keep the scope to the one task from step one.
Letting the vendor set the success metric. A metric chosen by whoever is selling the tool tends to flatter that tool. Use the metric your firm already tracked before the pilot started, from step four, not a new one the vendor proposes once results start coming in.
Treating the ninety days as a trial subscription, not a measurement. A pilot that nobody is actively measuring is just an early renewal decision wearing a different name. If there is no written before-number and no scheduled day-90 review, it is not a pilot in the sense this guide means.
Quietly extending past day 90 without a decision. The most common failure is not a bad result, it is no result at all: the pilot period ends, everyone is busy, and the tool keeps running by default. Put the day-90 review on a calendar in week one, before the tool is even switched on.
Skipping the downside number from step nine. A pilot that only tracks the upside cannot tell a good result from a lucky one. Write down what a failed pilot costs before you know whether it failed.
Step four asks for a before-number. Here is what that number is for, using a hypothetical for illustration only, not a published benchmark.
Scenario A. A firm times a document-intake task at 6 hours a week done manually. The tool cuts that to 2 hours a week — a 4-hour weekly saving. At a loaded rate of $75/hour across 48 working weeks, that is $14,400 a year in redeployed time. Training three people for 12 hours each, plus a 3-month pilot licence at $400/month, puts the downside from step nine at $3,900. Net first-year position: $10,500 — a clear case to extend into the full adoption plan.
Scenario B. The same tool, on a task that only saves 1 hour a week, returns $3,600 a year against the same $3,900 downside — a net loss of $300 before any licence renewal. Same tool, same firm, same pilot structure. The only thing that changed is the hours saved, and it flips the day-90 decision from adopt to kill.
Neither number is a claim about what your firm will see. Statistics Canada's Survey of Advanced Technology collects exactly this kind of adoption data for most sectors of the economy and explicitly excludes construction from its scope, which is why the method matters more than any number you read elsewhere — see measuring return on an AI rollout for the full method.
Step seven's barrier list is not evenly distributed. Where your firm sits changes which barrier is actually the one to watch.
None of these figures are construction-specific, which is exactly the barrier-list discipline step seven already asks for: name what you actually know, and do not borrow a number from a different industry to fill the gap.
No. The pilot structure in this guide assumes a configured, off-the-shelf tool and an accountable owner who is not a technical specialist. A pilot that genuinely needs custom development work is the exception covered in step ten, not the default case.
Treat that as a finding, not an inconvenience. A vendor unwilling to show you what its own tool is doing during a trial is telling you something about how the relationship will run after you sign a longer contract.
Generally, no, at a firm this size. Step two's single accountable owner cannot give two ninety-day pilots the same attention, and a mixed result across two pilots makes the day-90 decision in step eight harder to read cleanly. Run them in sequence.
A 30-minute call is enough to tell you whether it is worth building.