Treadstone Associates
Article · Compliance & records

Where do AI tools store your client data?

In four places, not one: your own workspace, the inference path, the logs, and the people with access. Vendors answer confidently about the second and vaguely about the third.

Treadstone Associates · Updated 2026

Key takeaways

  • • Most client data in a business AI tool never leaves your own tenant — the model reads it and the record of the interaction is stored with your other content. Microsoft documents exactly that for Microsoft 365 Copilot.
  • • Prompts and responses are their own class of record. The Office of the Privacy Commissioner asks that, unless otherwise required, prompts not be retained, used for secondary purposes or disclosed.
  • • Inference routing is not the same as storage. Microsoft states that Copilot calls to the model are routed to the closest datacentres in the region but can call into other regions when capacity is short.
  • • Sending information to a processor is a use by your firm, not a disclosure — and the OPC states that responsibility for protecting it stays with you.

The short answer

There are four answers, and treating them as one is where firms get into trouble.

  • Your own workspace or tenant. For a business tool wired into the systems you already run, the client data mostly sits where it already sat. What is new is the record of the interaction — the prompt, the response, and the reference to whatever the model read.
  • The inference path. The route a prompt takes to reach a model and come back. This is transient by design, but transient is not the same as local, and it is the part of the answer most likely to cross a border.
  • Logs and activity history. Everything kept for troubleshooting, abuse detection, billing and audit. This is where vendor answers get least specific, and it is the layer most likely to outlive your retention schedule.
  • People. Support engineers, subprocessors and anyone with an administrative role. A tool is only as confidential as the smallest set of humans who can read what goes through it.

What a well-documented answer looks like

Microsoft’s documentation for Microsoft 365 Copilot is a useful benchmark, not because the product is the right choice for every firm, but because the answers are written down and specific. It states that prompts, responses and data accessed through Microsoft Graph are not used to train the foundation large language models; that Copilot accesses content and context through Graph and can generate responses anchored in organisational data such as documents, email, calendar, chats, meetings and contacts; that interaction history from Copilot Chat and Teams meetings is processed and stored in alignment with the contractual commitments covering your organisation’s other Microsoft 365 content, encrypted at rest and not used for training; that administrators can view and manage that stored data through Content search or Purview, and can apply retention policies to it; and that users can delete their own activity history from the My Account portal.

On routing it is equally direct: calls to the model are routed to the closest datacentres in the region, but can call into other regions where capacity is available during high utilisation periods. That sentence is the honest version of what almost every hosted AI service does, and it is the reason “our data stays in Canada” is a claim to test rather than assume.

The layer vendors describe least well

Logs are where the gap usually is. Ask three questions and require the answer in the agreement rather than in a sales call: what is written to a log when a prompt is processed, how long is it kept, and who can read it. If the answer is that prompts are logged for abuse monitoring for a fixed period and readable by a named support function, that is a workable answer. If the answer is a shrug, treat it as a retention period of “indefinite” and decide accordingly.

The Commissioner’s principles for generative AI are unusually concrete here. Organisations should establish and abide by appropriate retention schedules for personal information including that contained within training data, system prompts and outputs, and those schedules should both limit retention where information is no longer required and ensure information is kept long enough for individuals to exercise access rights. The same document says that unless otherwise required, prompts should not be retained, used for secondary purposes or disclosed — which is a statement about the vendor’s log as much as about yours.

Sending it to a processor is a use, not a disclosure

This distinction decides who is answerable. The Office of the Privacy Commissioner’s guidelines for processing personal data across borders state that a transfer is a use by the organisation and is not to be confused with a disclosure; that information transferred for processing may only be used for the purposes for which it was originally collected; and that Principle 4.1.3 of Schedule 1 to PIPEDA requires organisations to use contractual or other means to provide a comparable level of protection while the information is being processed by a third party. They add that the organisation should have the right to audit and inspect how the third party handles and stores personal information, and should exercise it when warranted.

Treadstone Law’s answer on sharing customer information with a third-party provider covers the consent side of the same question, and its note on the definition of personal information is worth reading before you conclude that what you are sending is not personal information at all.

The five questions worth putting in writing

  • • In which country or countries is data at rest for this service, and does that commitment cover the model call as well as the storage?
  • • What is written to a log when a prompt is processed, for how long, and who can read it?
  • • Are prompts or outputs used to improve, tune or train any model, including a model dedicated to us?
  • • Which subprocessors are involved, and how are we told when that list changes?
  • • On termination, what is deleted, when, and what evidence of deletion do we get?

Worked example (illustrative)

A Manitoba firm evaluates two drafting tools. Tool A publishes documentation covering storage location, training use, log retention and administrative access; the commitments are repeated in the agreement. Tool B has a marketing page saying “enterprise-grade security” and a sales engineer who is confident but has nothing in writing.

The firm does not conclude that Tool B is unsafe. It concludes that Tool B cannot be assessed, which is a different and more useful finding, and it records the reason. The countable outcome is the proportion of tools in use for which the firm can answer all five questions from a document rather than from memory. That number is measurable today and was almost certainly not one hundred per cent before anyone asked.

Where this sits in the firm

This page is written for a firm that delivers work to a book of clients. If the question is really about the front desk — intake, scheduling, recall, reminders — that lives on the professional practice owners page. If it is about your own month-end, reconciliation and payables rather than client deliverables, that is accounting automation. The two overlap on tooling and almost never on risk.

Questions we get asked

If we use a tool inside our own Microsoft or Google tenant, has anything left?
The stored record generally has not, but the model call is a separate question — see the routing language above. Ask about both.

Does deleting the chat delete everything?
Not necessarily. Microsoft documents a user-facing deletion of Copilot activity history alongside administrative retention policies applied by the organisation, which are different controls with different effects. Assume the same split elsewhere and ask which one you are being shown.

Is a transcript of a client meeting different?
Operationally, no — it lives in the same four places. Practically, yes, because it is a verbatim record of a conversation rather than a derived summary, and it is far more likely to be sought later.

Where does the location rule actually come from?
Not from PIPEDA, which does not impose one. The rules that do are tax and records rules, covered in keeping client data in Canada.

Get a written answer from every tool you already run.

A 30-minute call is enough to tell you whether AI pays for itself here.