"Summarise this" is the instruction that produces the least useful output. Name the shape you need, and anchor every line to its source.
Key takeaways
You summarise documents in a client file by deciding what the summary is for, asking for that specific structure, and requiring a source reference on every line. A chronology, an obligations register and an issues list are three different products; asking for "a summary" gets you a fourth thing that is none of them.
The second half matters as much as the first. A summary with no anchors cannot be checked, so a reviewer has to read the file anyway — which means the summary saved nothing.
A chronology. One row per dated event: date, what happened, who did it, source document and page. This is the highest-value output on almost any file with a dispute or a sequence in it, and the easiest to check, because dates are objective.
An obligations register. One row per commitment: who owes what, to whom, by when, under which clause. Built from contracts, variations and correspondence, it is what an engagement actually needs when the question is "what are we on the hook for".
An issues list. One row per open question, with the competing positions and the documents supporting each. The discipline here is that the model must not choose between them.
A decision log. What was decided, by whom, on what date, on what information, and what was expressly not decided. Firms rarely keep one, and it is the artefact they most want to have two years later.
Require a citation format and enforce it — document name and page, at a minimum. Then spot-check: pick five lines at random, open the pages, confirm. Do this every time on a new document type and periodically thereafter, because it is the only thing that measures whether the summary is true rather than plausible.
Anchoring also solves the argument about whether the tool "hallucinated". A line with no anchor is not a claim you have to litigate; it is a line that fails the format and gets removed before anyone reads it.
A consultancy takes over a stalled programme. The file is a signed agreement, two change orders, eleven monthly status reports, a steering committee pack, and roughly 200 pages of email.
The instruction is not "summarise the file". It is three separate runs. First a chronology from the status reports and steering packs only, one row per reported milestone with its reported status and the page it came from. Second an obligations register from the agreement and change orders, one row per deliverable with the clause reference. Third an issues list from the correspondence, with each position attributed.
The chronology immediately shows a milestone reported green in three consecutive months and then red, with no intervening change order — a contradiction between documents the model was told to surface rather than reconcile. That single row is worth more than any prose summary of the file, and a person found it in a minute because the chronology put the two sources side by side.
Version confusion. Three drafts of the same agreement in one folder, and the summary quietly blends them. Fix it by summarising per document and reconciling afterwards, never by pointing the tool at a folder.
Length collapse. Long files get summarised unevenly, with early pages over-represented. Chunking by document rather than by page count keeps the weighting honest.
Confident resolution. Asked to summarise contradictory material, a model tends to produce one coherent story. That is the opposite of what a professional file needs. Instruct it to report disagreement as disagreement.
Quantities and dates. These are where a fluent summary is most often wrong and least often questioned. Treat every figure in a summary as unverified until someone has opened the page.
A client file is personal information the moment it contains identifiable individuals, and the Office of the Privacy Commissioner is explicit that generative AI tools do not sit outside Canada's existing privacy frameworks — its principles for responsible, trustworthy and privacy-protective generative AI are written for organisations using such systems as well as those building them.
The baseline obligations sit in the OPC's summary of PIPEDA requirements, and whether PIPEDA or a provincial statute applies to your firm is worth settling once rather than per engagement — see does PIPEDA apply to your business and what counts as personal information.
Then there is the client's own confidentiality expectation, which is usually contractual as well as professional. If your engagement terms or an NDA restrict disclosure, uploading the file to a third-party service is a disclosure — see breach of a confidentiality clause, and settle vendor terms in advance, which is the same discipline as any cybersecurity and data privacy due diligence exercise.
The summary is an input to professional work, never the work. Nobody advises from a summary, and nobody signs one. What a person does is decide which lines matter, open the pages behind those lines, and form a view — which is exactly the part the reader is paying for.
If the file is a legal one, the risks are different and sharper; that is dealt with separately in AI for summarising legal documents.
How long should the summary be?
Long enough to be checkable. A one-paragraph summary of a 300-page file is a comfort object. A structured table with 60 anchored rows is a working document, and it is faster to use.
Can we send the summary to the client?
Only after a person has verified it, and only if the engagement contemplates it. A summary is a draft, and a draft that leaves the building has become a deliverable.
Does it help on files we already know?
Often more than on new ones, because you can spot-check it instantly and because a chronology of a familiar file surfaces the gaps that familiarity hides.
A 30-minute call is enough to tell you whether AI pays for itself here.