Treadstone Associates
Article · 8 min read

Structured vs unstructured data, explained

Every AI project runs into the same two words within the first week: structured and unstructured. The distinction sounds technical, but it decides what an AI tool can actually do with a business’s data on day one, and what has to be extracted, cleaned, or redacted before it can do anything at all.

Treadstone Associates · Updated 2026

Key takeaways

  • • Structured data lives in fixed fields a database can query directly — a CRM record, an invoice line, a closing date. Unstructured data — email, PDFs, call recordings, contracts — carries no such schema.
  • • In Canada’s own adoption data, the two are used almost equally: data analytics and text analytics are the two most-reported AI applications among Canadian businesses that use AI at all.
  • • Personal information does not respect the line between the two. Canada’s federal privacy statute applies to a client’s name in a spreadsheet cell exactly as it applies to the same name buried inside a paragraph of an email.
  • • Redacting a structured field is a delete key. Redacting the same information once it is embedded in free text or audio is a genuinely harder, more error-prone task — which is why unstructured data carries more practical privacy risk even though the legal rule does not change with format.

A business that has never built anything with AI tends to use “data” as if it were one substance. It is not. The single most useful distinction for deciding what an AI tool can touch first — and what needs work before it can touch anything — is whether the data is structured or unstructured.

What “structured” actually means

Structured data sits in a defined schema: named fields, consistent types, rows and columns a database engine can filter, sum, or join without guessing. A client’s file status in a CRM, a mortgage amount in a loan-origination system, a closing date in a spreadsheet column — each is structured because the system already knows what kind of thing it is looking at before a single record is read. This is the layer most AI tools reach first, not because it is more important, but because it requires the least additional work: the meaning of the data is already encoded in where it sits.

What “unstructured” covers — and why it is most of what a business actually has

Unstructured data is everything without that schema: emails, scanned contracts, meeting notes, call recordings, chat transcripts, PDFs. Most of what accumulates inside a Canadian business over years of operating is this kind of material, and none of it tells a system what it is until something — a person or a model — reads and interprets it. The U.S. National Institute of Standards and Technology’s AI Risk Management Framework, cited by Canada’s privacy regulators as the reference point for evaluating an AI tool’s validity and reliability, notes that “maintaining the provenance of training data and supporting attribution of the AI system’s decisions to subsets of training data can assist with both transparency and accountability” — and adds, pointedly, that “training data may also be subject to copyright and should follow applicable intellectual property rights laws.” NIST AI RMF, AI Risks and Trustworthiness That warning matters far more for unstructured material — contracts, articles, correspondence — than for a numeric database column, because that is exactly where someone else’s copyrighted work tends to live. Canada has no exception in the Copyright Act for using that material to train a model; fair dealing under section 29 covers “research, private study, education, parody or satire” and nothing broader. Copyright Act, s.29

Why AI treats the two differently, mechanically

Structured data can usually be handed to an analytics-oriented AI tool close to as-is: the fields are already labelled, so the tool’s job is pattern-finding, not interpretation. Unstructured data has to be converted into something a model can actually use first — indexed for retrieval, transcribed, or used to fine-tune — and in every case a person has to decide what gets extracted and what gets discarded. That decision is where quality, bias, and privacy risk all get introduced at once, long before a model produces a single output.

The Canadian split in practice

Statistics Canada’s second-quarter-2026 survey of business conditions found that among the 19.2% of Canadian businesses using AI — a proportion that “has tripled since the second quarter of 2024 (6.1%)” — the most commonly used application was data analytics at 36.6%, followed closely by text analytics at 34.5% and virtual agents or chat bots at 28.2%. Statistics Canada, Q2 2026 AI use survey That is a near-even split between the structured layer (data analytics) and the unstructured layer (text analytics), which cuts against the assumption that businesses are only reaching for the easy, already-tabular data. In practice, Canadian businesses are working both layers roughly in parallel.

Personal information does not care about the difference

Canada’s privacy commissioners’ joint guidance on generative AI advises that organizations should “use anonymized, synthetic, or de-identified data rather than personal information where the latter is not required to fulfill the identified appropriate purpose(s).” OPC, generative AI principles Nothing in that principle turns on whether the personal information sits in a database column or a paragraph of prose — the obligations under the federal Personal Information Protection and Electronic Documents Act attach to the information itself, including once it has been handed to a third party for processing. PIPEDA, Schedule 1, cl.4.1.3 In a structured record, removing a client’s SIN or account number is a matter of dropping a column. In an unstructured record — the same number typed into an email, or spoken aloud on a recorded intake call — there is no column to drop. The information has to be found first, and finding it reliably inside free text or audio is a materially harder task than filtering a query. That gap, not the law itself, is why unstructured data is where most AI privacy incidents actually originate.

A worked example

Consider a Canadian mortgage brokerage. Its CRM holds structured records: client name, file status, mortgage amount, closing date — each field already labelled, already queryable. Its shared inbox and recorded intake calls hold the unstructured version of the same relationship: the same clients’ income, debts, and life circumstances, described in free-flowing sentences and audio rather than fields. An AI tool can safely summarize the CRM’s file-status trends today with modest governance work, because the fields are already known quantities. Pointing the same tool at the inbox or the call recordings without first identifying and handling the personal information embedded in them is a materially different, and materially riskier, undertaking — not because the law changes, but because nothing in an email thread announces where the sensitive part starts and ends.

Related: what counts as good training data, why business data is rarely AI-ready, and how AI handles messy business data.

Common questions

Does “unstructured” mean a document is unusable by AI?

No. It means the document needs a processing step — extraction, indexing, or transcription — before a model can use it reliably. Text analytics is the second most common AI application among Canadian businesses precisely because tools now do that processing step routinely.

Is a PDF structured or unstructured data?

Usually unstructured, even when it looks organized to a human reader. A PDF invoice has a visual layout but not a machine-readable schema, so a model generally has to interpret its layout rather than read defined fields, unless a structured version of the same data exists elsewhere in the business’s systems.

Does Canadian privacy law treat structured data more leniently than unstructured data?

No. PIPEDA’s Schedule 1 obligations attach to personal information itself, regardless of the format it is stored in. The practical difference is that finding and protecting that information is far easier in a labelled field than inside free text or audio.

Not sure which of your data is actually usable yet?

A short call is enough to map what is already structured, what needs work, and what should not be touched without governance first.