Treadstone Associates
Article · Client onboarding & engagement lifecycle

AI for the new-client information form

The reason you keep retyping client details is that the firm has no single client record — only a series of forms. Fix the record and the retyping disappears.

Treadstone Associates · Updated 2026

Key takeaways

  • • Separate the standing client record from the per-engagement delta and most duplication vanishes.
  • • Extraction is reliable for structured documents and weak on anything handwritten or reformatted.
  • • Collect less: the identified purpose sets the limit, not what might be useful one day.
  • • Both the Income Tax Act and the Excise Tax Act impose six-year retention on the underlying records.

The short answer

Stop treating it as a form. A practice serving a book of clients has two different things tangled together: a standing record that is true about the client until it changes, and an engagement delta that is specific to this year’s work. Most firms re-collect both every time, because the only place the standing record exists is inside last year’s email thread.

Once the two are separated, AI has an obvious job: read what the client has already sent and populate the standing record, so the request that goes out asks only for what is genuinely missing.

Why the same information gets asked three times

Three causes, in order of frequency. The information lives in a document rather than a field, so nobody can query it. The information was captured by a person who has since left. And the request template is service-specific, so the corporate team asks for the incorporation details the tax team already holds.

None of these is an AI problem. They are a data-model problem that AI makes cheaper to fix, because the historical documents can be read at scale instead of re-keyed.

The standing record

Small, stable, and worth getting right once. Typically: legal name and any operating names, business number and programme accounts, incorporation jurisdiction and date, fiscal year end, directors and signing authority, principal contacts with their roles, banking institution, and the firm’s own file references. It changes a few times in a decade.

The engagement delta is everything else: this year’s trial balance, the new lease, the shareholder change, the slips. It is large, it is transient, and it should never be the thing that carries the client’s standing details.

What extraction does well, and badly

  • Reliable: structured documents with consistent layouts — registry printouts, articles, standard statements, prior-year returns produced by known software.
  • Workable with review: letters and agreements, where the fields exist but the phrasing varies.
  • Weak: handwriting, photographs of documents taken at an angle, and anything a client has retyped into their own template.
  • Never unreviewed: anything that becomes a figure in a filing. Extraction produces a candidate value with a source; a person confirms it.

A useful rule for a practice: extraction may populate a field, but the field stays marked unconfirmed until somebody has looked at the source. Firms that skip that step discover the cost during a review, at the worst possible moment.

Collect less, not more

The temptation with an automated request is to ask for everything, because asking is now free. Resist it. The Privacy Commissioner’s generative AI principles restate the familiar rule — limit collection, use and disclosure to what is needed for the explicitly specified, appropriate purpose — and add an openness expectation about what happens to the information at each stage. A request list that has grown by accretion over eight years is a liability, not thoroughness.

Retention is the same principle running forward instead of back. PIPEDA’s own Limiting Use, Disclosure, and Retention principle says personal information is to be retained only as long as necessary to fulfil the purpose it was collected for, and destroyed, erased or anonymised once that purpose is served. That sits in real tension with the tax retention rules below, which set a floor, not a ceiling: a standing-record field you collect on spec, on the theory it might be useful next year, is exactly what this principle says not to keep. Treadstone Law’s guide to PIPEDA customer data rules covers what counts as a legitimate purpose for a small practice.

If your practice holds personal information about clients’ employees or customers as a by-product, the privacy policy requirements for an Ontario small business are the starting point for what you must be able to say about it.

Retention: what the law actually requires

The request list should be traceable to why the record has to exist at all. Under section 230 of the Income Tax Act, every person carrying on business must keep records and books of account in such form and containing such information as will enable the taxes payable to be determined, and must retain them until six years from the end of the last taxation year to which they relate; records kept electronically must be retained in an electronically readable format for that period.

The GST/HST side is parallel: section 286 of the Excise Tax Act requires records adequate to determine liabilities and obligations, retained for six years after the end of the year to which they relate, and — a detail firms miss — kept in Canada in English or in French unless the Minister authorises otherwise. On the corporate side, Ontario corporate record retention has its own schedule.

Where the software sits

Practice management platforms are where this belongs, because the standing record has to be adjacent to the work. Karbon’s documentation describes AI features built into the platform that draft client emails, summarise long threads, flag missed time entries and suggest team assignments, drawing on the firm’s own data across clients, jobs and communication, with an assistant it calls Kai described as an early-access extension of that. Treat those as capability descriptions to test against your own files — the question that decides it is whether the standing record ends up in a queryable field or in another thread.

A worked example

A practice with about 180 corporate clients spends the first three weeks of every busy season re-collecting details it already holds. It does one thing differently: it defines the standing record as fourteen fields, then runs last year’s permanent files through extraction to populate them, with a junior confirming each field against the source document.

The output is a short list of clients where something is genuinely missing or has changed. Those clients get a specific request naming the two items outstanding, instead of a generic eleven-item form. Response rates improve for an unglamorous reason: a request that asks for two things you can find in a drawer gets answered, and one that asks for eleven gets postponed.

The following year the same run takes an afternoon, because the standing record already exists and only the delta is in play. That is the whole return on the exercise, and it compounds.

Questions we get asked

Can AI fill in the client’s form for them?
It can pre-populate what you already hold so the client confirms rather than composes, which is the single biggest lever on response rate. It should not invent a value it cannot source from a document.

What about the information we collect for identity checks?
That is a separate regime with its own listed methods and record-keeping rules. Read does AI satisfy client identification rules? before folding those fields into a general intake form.

Do we need to keep the source documents once the fields are populated?
Yes. The statutory retention obligations attach to the records, not to your database. Section 230 and section 286 both set six years, and a populated field is not a substitute for the document behind it.

See where AI pays off first in your practice.

A 30-minute call is enough to tell you whether AI pays for itself here.