Treadstone Associates
Article · 10 min read

How do you calculate ROI on an AI project?

Pick one repeated, document-shaped task, measure its cycle time and error rate before you buy anything, then value the change in whatever the business actually gets paid for. Return is the difference between two measurements of the same process — so if you never took the first one, you have no business case, only the vendor’s. Here is the arithmetic, and the Canadian reason cycle time is worth more in construction than in most sectors.

Treadstone Associates · Updated 2026

Key takeaways

  • • The denominator is licence + unrecoverable tax + implementation + permanent review time + exit, not the licence alone.
  • • Count hours only where the hour is redeployed into something the business is paid for.
  • • Cycle time is the underrated benefit: prompt payment deadlines in Ontario run off documents, so earlier paperwork moves cash.
  • • Take a two-week baseline before signing, and repeat the same measurement after eight weeks of live use.
  • • Never model the vendor’s figure — a performance claim must rest on an adequate and proper test, so ask what the test was.

Pick one repeated, document-shaped task; measure how long it takes and how often it is wrong before you buy anything; then value the change in whatever the business actually gets paid for. Return on an AI project is not a property of the software. It is the difference between two measurements of the same process, and if you never took the first measurement you cannot compute it.

That is why most AI business cases in construction fail on arithmetic rather than technology: nobody wrote down the baseline, so the only figure available afterwards is the vendor’s.

The shape of the calculation

Return is the annual value of the change divided by the annual all-in cost. The denominator is the part people get wrong, because the licence is only part of it:

  • Licence and unrecoverable tax.
  • Implementation — getting documents, templates and price data into usable shape.
  • Review time, every week, forever. Someone competent reads the output.
  • Migration and exit — what it costs to get your records back out.

The full cost side is set out in what AI actually costs a small contractor. This article is about the numerator.

Three kinds of benefit, and which ones are real

Hours avoided. Real, but only if the hour goes somewhere. An hour saved by a salaried estimator who now bids one more job is revenue. An hour saved that becomes an hour of something unmeasured is not a return, it is a working-conditions improvement — worth having, but do not put it in the business case as money.

Cycle time. Often the biggest and most overlooked, because in Canadian construction the paperwork clock and the money clock are connected. Ontario’s prompt payment and adjudication regime under the Construction Act runs on documents, and interim adjudication is administered by ODACC. Our sister firm’s explanation of the prompt payment rules and the timelines that follow each document sets out which document starts which clock — we do not restate the day counts here, because they belong to the statute and not to a blog. The point for a business case is simply that the clocks are document-driven, so issuing the document earlier moves the payment earlier, and that is a change your bank statement can confirm.

Errors avoided. Real but hard to attribute. Count only errors that would have cost money and that you can evidence — a missed lien deadline, a mispriced allowance, a change order that went unbilled. Our sister firm’s note on how long a contractor has to register a lien after finishing work is a useful reminder of which dates are expensive to miss.

Take the baseline before you sign

Two weeks is usually enough. For the chosen task, record: how many of them per week; minutes from trigger to finished document; how many required rework; how many went out late; and who touched each one. Keep it in a spreadsheet, not in anyone’s head.

Do the same measurement again after eight weeks of live use. The comparison is the return. Anything else — a percentage from a case study, a vendor’s benchmark, an industry average — is a claim about someone else’s business.

Do not put the vendor’s number in your model

A performance claim that is not based on an adequate and proper test, with the burden of proof on the person making the representation is reviewable conduct under the Competition Act, and the Competition Bureau’s own guidance to advertisers is to make performance claims only if they have been substantiated by an adequate and proper test. That is a rule about what a vendor may say to you. It is also a good filter: ask what the test was, on what document set, judged by whom. If the answer is a customer anecdote, treat the figure as marketing and use your own baseline instead.

The exposure for making an unsubstantiated performance claim is not abstract: under s. 74.1(1)(c), a court can order a corporation to pay an administrative monetary penalty of up to $10,000,000 for a first order ($15,000,000 for a subsequent one), or three times the benefit derived from the deceptive conduct if that amount can be determined. An order like that isn’t a one-time cost, either — it applies for up to ten years unless the court sets a shorter period (s. 74.1(2)).

A worked example

The figures below are illustrative of the method, not a benchmark — the whole point is that the numbers have to be yours. A nine-person general contractor in Ontario picks one task: turning field notes and photographs into the daily report and the weekly owner update. Its own baseline over two weeks: 5 reports a week, averaging 38 minutes each from site walk to sent, 1 in 5 needing a correction after issue, and 2 of 10 issued the following day rather than the same day.

After eight weeks with an extraction tool and a fixed template: the drafting is faster, but the project coordinator now spends 8 minutes reviewing each draft, and that review time is permanent. The honest numerator is (old minutes − new minutes − review minutes) × reports per year × the coordinator’s internal rate, plus the value of the two same-day issuances that now happen on time, plus any corrections avoided.

The firm also finds something the spreadsheet did not predict: because the owner update now goes out on Friday rather than Monday, one contentious change order is being discussed while the work is fresh instead of a month later. That is real, and it belongs in the memo even though it does not belong in the ratio.

What to write down before you buy

  • Volume per week, and the seasonal peak — per-document pricing punishes tender season.
  • Minutes from trigger to finished document, measured, not estimated.
  • Rework rate and late rate, with the reason for each.
  • Who signs. If the answer is “nobody yet”, fix that before the pilot, not after.

Common questions

How long before we know?

Long enough to cover a full cycle of the task, including its worst week. A tool evaluated only in a quiet month will look better than it is. Eight weeks of live use against a two-week baseline is a defensible minimum for a document-shaped task.

The tool is right most of the time. Whose job is the rest?

Yours. BCFSA puts it plainly for licensees: AI tools are designed to guess at answers they do not have complete information for, and you could be held accountable for any inaccuracies propagated by the AI. CREA’s AI FAQ says the same in its own terms — the use of AI does not alter your obligations, and AI-generated information should be independently verified before it is relied on or shared. Every business case must therefore carry the review cost; a model that omits it is not optimistic, it is wrong.

Does the calculation change if we build our own?

Yes, substantially — the denominator grows to include development and the governance work that comes with it. See should you build AI or buy off-the-shelf?. And if the alternative you are really weighing is a person rather than a product, read hire another admin, or automate the work?.

See where AI pays off first in your business.

A 30-minute call is enough to tell you whether AI pays for itself here.