Treadstone Associates
Guide

Measuring return on an AI rollout

No one has published a Canadian number for what an AI rollout returns in construction. That is not a gap in this guide — it is the honest starting point for building your own.

Treadstone Associates · Updated 2026

Key takeaways

  • • No published Canadian figure exists for AI return on a construction rollout — measure your own, from records you already keep.
  • • Record a real before-number for at least one full cycle before judging anything.
  • • Separate the tool's effect from ordinary market noise; StatCan's own monthly figures move by material amounts without any AI involved.
  • • Training time is a genuine, measurable cost. Track your own hours the same way StatCan tracks the national figure.

STEP 01 OF 10

Start from the honest baseline

StatCan found that 66.7% of Canadian businesses had no plans to adopt AI over the following twelve months, and 78.1% of that group said it was simply not relevant to what they do. That is the real state of the market: most Canadian businesses, construction included, have not built this measurement yet, because most have not adopted the tool that would require it. You are not late to a proven trade with a known return. You are building the first internal measurement your firm has ever had a reason to build. Treat it that way — slowly, and from your own numbers.

STEP 02 OF 10

Pick metrics your firm already logs

Days-to-quote, rework hours, bid win rate, days-to-invoice, RFI turnaround — whatever already lives in your job-costing or accounting system. Do not adopt a metric because a vendor suggested it; a vendor-suggested metric is usually chosen because it flatters that vendor's tool, not because your firm can actually verify it independently. This is the same before- number discipline described in the ninety-day pilot guide, step four — the pilot and the measurement are one exercise, not two.

Pick two metrics at most for a first rollout, not five. A long scorecard invites picking whichever line happened to improve and calling that the result; two metrics, agreed in advance, are harder to cherry-pick from.

STEP 03 OF 10

Record a full cycle before switching anything on

The federal privacy commissioners' own guidance frames evaluating a tool as an evidence-based question: is it "necessary and likely to be effective" for the stated purpose. You cannot show "likely to be effective" without a real before-number, and one week of data is not a cycle — a slow month followed immediately by a busy one will make almost any tool look transformative by accident. Record at least one full quoting or billing cycle before you compare anything.

STEP 04 OF 10

Separate what the tool touched from everything else that changed

Material prices, staffing changes, and a large contract landing or falling through can move your numbers as much as any tool. StatCan’s release on investment in building construction for May 2026 recorded the total value edging down 0.3% to $23.4 billion for the month, while growing 5.9% year over year — and provincial swings inside that single month ran from a $52.5-million increase in British Columbia's multi-unit segment to an $80.1-million decrease in Alberta's. That is ordinary market movement, in a single month, with no AI tool involved anywhere in it. Before crediting or blaming a rollout for a change in your own numbers, ask what else moved that quarter.

STEP 05 OF 10

Track the training line, because it is a real cost

StatCan found that 44.4% of AI-using Canadian businesses changed training or staffing practices, and among firms with 100 or more employees, 30.2% used external consultants or vendors for it — a cost that smaller firms mostly cannot outsource and instead carry as staff time. Log your own hours here rather than assuming a national percentage applies to your firm; the point of the national figure is only that this cost is real and typical, not that it matches your size or trade.

STEP 06 OF 10

Watch for the change nobody put in the business case

A StatCan study on Canadian employment trends through the recent AI era carries its own caveat, worth borrowing directly: "It is unclear whether more recent trends reflect the advent of AI, other economic factors such as labour market adjustments after the COVID-19 pandemic, rapid demographic shifts, recent trade tensions with the United States or a combination of factors." Apply the same humility to your own quarter. If win rate improved the same month a strong estimator joined the team, the estimator is at least as likely an explanation as the tool.

STEP 07 OF 10

Write the number down even if it disappoints

Feed the day-90 or day-180 result back into the adoption plan, whatever it says. A null or disappointing result, written down with the numbers behind it, stops the firm re-litigating the same pilot from scratch a year later at the same cost, having learned nothing new the second time either.

This is also the record that protects the person who championed the tool. A written, honest result — even a modest one — is a far stronger position than an undocumented impression that gets relitigated from memory at the next budget meeting.

STEP 08 OF 10

Never repeat a vendor's stated figure as your own result

A vendor's claimed accuracy rate, time saving, or return on investment is a marketing claim about their product in general, not a measurement of what happened at your firm. As of the most recent Canadian releases used throughout this hub, Statistics Canada publishes adoption, application, and barrier data for AI use in Canadian businesses — not a return figure, for construction or any other industry. None exists to borrow. Measure your own, and only report your own.

If a vendor's sales material cites a study, ask which country it was measured in and on what kind of work. A benchmark from a different industry, a different country, or a different-sized firm answers a different question than the one your own before-and-after numbers answer.

STEP 09 OF 10

Understand why the number you're looking for doesn't exist yet

Statistics Canada's Survey of Advanced Technology, its main federal instrument for technology-adoption data, states its own scope plainly: it collects data from "Canadian businesses with more than 10 employees and sales of over $250,000 in all sectors of activity except construction." Construction is not an oversight in one release — it is written out of the survey's design. That is a structural reason no Canadian body publishes a construction AI-ROI figure, not a gap someone forgot to fill.

Knowing this changes what you should do with a number you find somewhere else. A consultant's figure, a vendor's case study, or an economy-wide statistic from a survey that explicitly excludes your industry is not a benchmark you are falling short of or exceeding — it is a number from a different population, and the method in this guide exists because that population does not include you.

STEP 10 OF 10

Read a cross-sector figure correctly before you use it at all

BDC's own research found that among Canadian SMEs, businesses using AI in 2025 "were 24% more productive than those that didn't". That figure is real, it is Canadian, and it is still the wrong number to cite as construction's expected return: it is a cross-sector correlation, not a construction-specific measurement, and it says nothing about which of those businesses adopted AI because they were already better-run rather than becoming better-run because of it.

The honest use of a figure like this is as motivation to build your own number, per step one, not as a substitute for one. If a figure did not come from your own before-and-after records on your own task, it cannot answer the question this guide's title asks.

Common mistakes

Measuring too soon. A number pulled two weeks after go-live mostly measures the learning curve, not the tool. Wait for at least the one full cycle described in step three before drawing any conclusion.

Comparing a good month to a bad month. Picking the single best week after rollout and the single worst week before it, even unintentionally, will always show an improvement. Compare full, equivalent periods, not cherry-picked ones.

Treating a null result as a personal failure. A pilot that shows no measurable benefit did its job: it answered the question honestly, using real numbers, before the firm spent a year committed to something that was not working. That is a successful measurement, not a failed one.

Letting enthusiasm substitute for the number. "Everyone likes using it" is a real signal worth noting, but it is not the same signal as the before-and-after metric from step two. Report both, and do not let the first stand in for the second.

Measuring the tool instead of the task. A tool can be technically accurate and still not move the metric your firm actually cares about, if it was never aimed at the step that was slow in the first place. Confirm the measured task matches the one the pilot was scoped to, not a nearby one that happened to be easier to instrument.

Stopping the measurement once the answer looks favourable. It is tempting to stop tracking the moment the early numbers look good, but a short favourable window is exactly the kind of result that fails to hold up over a full quarter. Keep recording through the full period set out in step three, not just until the first encouraging data point arrives.

Citing the Survey of Advanced Technology's headline figure for your own industry. The 60.6% adoption figure it publishes is real — and it is drawn from a survey that, by its own stated scope, does not include construction at all.

The same tool, measured two honest ways

Steps one through eight already lay out the method. Here is what following it, and not following it, produces on the same hypothetical rollout.

Scenario A. A firm records six weeks of manual baseline on a specific estimating task before the tool is switched on, per step three. The task drops from 5 hours to 3 hours a week across four estimators. That is 8 hours a week firm-wide, worth $600 a week at a $75 loaded rate, or $28,800 across 48 weeks — a number this firm can defend because every input is its own record.

Scenario B. The same firm instead applies BDC's 24% productivity figure to its estimating department's total payroll cost, producing a headline number with no relationship to what the tool actually changed in that department, because the 24% figure was never measured on estimating tasks, or on construction firms, at all.

Both numbers get called "return on the AI rollout." Only one of them was actually measured, and it is the one built from six weeks of the firm's own timing records, not the one borrowed from a national cross-sector average.

What actually gets published, and what doesn't

This guide's promise of no ROI figure is a deliberate choice, not an oversight. Here is exactly what the record does and doesn't support.

  • Published, and usable directly: economy-wide adoption rates by sector and firm size (StatCan, already cited across this cluster), and real financing terms such as BDC LIFT's $25,000 to $5 million range.
  • Published, but not a construction figure: BDC's 24% cross-sector productivity finding and the Survey of Advanced Technology's 60.6% economy-wide adoption rate — both real, neither about your industry.
  • Not published anywhere, by design of the survey: a construction-specific AI-ROI number, because the federal survey that would normally produce one explicitly excludes the industry.

The gap in the third row is exactly why steps one through eight exist. It is not a gap this guide is going to paper over with a number from the second row.

Frequently asked

If no Canadian figure exists, can we use a US or vendor-published ROI figure instead?

The same objection applies more strongly. A US figure carries a different regulatory and labour-cost environment on top of the cross-sector problem described in step ten; a vendor figure carries an obvious incentive problem on top of both. Use your own six-week baseline instead.

How precise does the baseline in step three need to be to count?

Precise enough that the person who did the timing would stand behind the number in a room, not precise to the minute. A rough but honestly-recorded baseline beats a polished number reconstructed from memory after the fact.

What if the honest number, from step seven, is a loss?

Then write it down as a loss. A firm that measures honestly and finds a negative return has learned something real about that specific tool and task; a firm that never measures at all has learned nothing, regardless of how the pilot felt to use.

See where this pays off first in your firm.

A 30-minute call is enough to tell you whether it is worth building.