Treadstone Associates
Article · 8 min read

Can AI cite its sources reliably?

A language model can produce a citation that looks exactly like a real one — a report title, a working-looking URL, a plausible case name — without having retrieved anything at all. Whether it is actually reliable depends on what sits behind the sentence, not on how convincing it sounds.

Treadstone Associates · Updated 2026

Key takeaways

  • • By default, a generative model predicts the next plausible token, including tokens shaped like citations, whether or not a source exists.
  • • Canada's Cyber Centre instruction on this is unambiguous: always validate sources, because outputs “can be incorrect.”
  • • A citation can also be genuine but outdated or narrow if the model's training data was time-bounded or drawn from a single source.
  • • Format is not evidence: a well-formatted citation and a fabricated one look identical until someone opens the link.

Why a citation can look real and be invented

A generative system, by default, produces its answer the same way regardless of whether the sentence is an opinion or a citation: it predicts the next plausible token based on patterns in its training data. A citation-shaped string — a report title, a section number, a URL format — is just another pattern the model has seen often enough to reproduce convincingly. Unless the system has been specifically connected to a retrieval step that actually looks something up at the moment it answers, producing a citation is a stylistic act, not a lookup.

This is precisely why a fabricated citation tends to be so plausible rather than obviously wrong. The model has seen thousands of real citations during training and has learned their shape extremely well — the punctuation, the structure of a section number, the cadence of a report title. It reproduces that shape faithfully even on the occasions when there is no real document behind it, because shape and substance are not the same thing it is tracking.

What Canada's own guidance says to do about it

The Canadian Centre for Cyber Security is direct about the underlying risk, warning plainly that generative AI outputs can be wrong, biased, or simply fail to account for relevant factors, and its instruction is not conditional: you should always be aware of and validate your sources to verify whether the content being presented is accurate. That instruction applies with equal force to a plain factual claim and to a citation dressed up to look authoritative — the format of the sentence does not change what is required of the reader.

It is worth being specific about what “validate” means in practice for a citation: not re-reading the sentence and deciding it sounds credible, but opening the actual link or searching for the actual report by name and confirming the quoted or paraphrased content is genuinely there. The first of those two habits feels like due diligence and accomplishes nothing; only the second one does.

A citation can be genuine and still misleading

Fabrication is not the only failure mode. Canada's federal-provincial privacy principles for generative AI note that providers should inform organizations using generative AI about any known issues or limitations about the accuracy of generative AI outputs. This may include where the training dataset is time bounded (i.e. only contains information up to a certain date); where it is from a single, non-representative source. A source that genuinely exists can still be a poor answer to your question if it is out of date or was the only source of its kind the model happened to see — a different problem from an invented citation, but one that the same discipline of checking the source, not just its existence, catches.

A real report from several years ago, cited to support a claim about a rule that has since changed, is arguably a worse outcome than an obviously fake citation: it survives a cursory check (the source exists) while still misleading the reader about the current state of the world. Opening the source has to include checking its date, not only confirming it is a genuine document.

What actually makes a citation reliable

A tool wired to retrieve a passage from an actual document store at the moment it answers can hand you a link to something you can open and check yourself — that is a fundamentally different guarantee than a model recalling what a citation typically looks like. fine-tuning vs prompting vs retrieval covers how retrieval, fine-tuning and prompting each change a model's behaviour differently, including which of them actually connects an answer to a checkable source.

The NIST framework's definition of accuracy is a useful test to apply here: it is the closeness of results of observations, computations, or estimates to the true values or the values accepted as being true — a definition that only makes sense if there is a true value to compare against. A citation you have not opened has not been compared to anything.

A worked example: the difference a single open tab makes

Suppose an employee asks a generative tool for the source behind a statistic they plan to put in a client memo, and the tool returns a report title, an author and a page number. If the employee copies that citation straight into the memo, the citation has cost the reader nothing to produce and has been checked by nobody. If the employee instead searches for the report by the name given, opens it, and either finds the exact figure on the page cited or does not, the citation has gone from an unverified string to an actual fact — and the small amount of time that search took is the entire difference between the two outcomes. Nothing about the tool changed between those two versions of the memo; only whether a person did the one step that makes a citation mean anything.

A practical habit for a business, not just an individual user

The same discipline that applies to one person checking one answer should be written into how a business uses these tools at all, not left to whoever happens to be careful that day. Any workflow that lets an AI-generated citation reach a client, a filing, or a decision without a person opening the link first is a process gap, not a technology limitation — the fix is a review step in the workflow, not a better model.

There is also a direct legal exposure, not just a reputational one, when an unverified figure reaches the public. Section 74.01(1)(a) of the Competition Act makes it reviewable conduct to make a representation to the public that is false or misleading in a material respect, and the same section reverses the burden of proof for a performance claim: the business making it, not a regulator, has to show the claim was based on an adequate and proper test. Sourcing a false figure from an AI tool rather than writing it yourself changes none of that exposure. treadstonelaw.ca’s overview of misleading advertising rules covers the general framework this sits inside.

See also common AI myths in business.

Common questions

Does a nicely formatted citation mean it is real?

No. Format and reality are unrelated in a generative system — it can produce a plausible-looking, correctly formatted citation for a source that does not exist, because formatting is just another pattern it has learned to reproduce, independent of whether the underlying document does.

Is this the same thing as the tool being biased?

No, they are different failure modes. Bias comes from what the training data over- or under-represents; a fabricated citation comes from the model completing a pattern with no retrieval step behind it. See common AI myths in business for the bias question specifically.

Should a business ever use an AI-generated citation without checking it?

No. The same rule applies every time: open the link, confirm the source says what the sentence claims, before it goes into anything a client, a regulator or a decision-maker will see.

Does it matter whether the model was asked to ‘be careful’ or ‘only cite real sources’?

An instruction like that can help at the margins, but it cannot give the model a retrieval step it does not have. Asking a system without retrieval to be careful about citations changes how carefully it tries to sound careful; it does not change whether a lookup actually happened.

Where this goes next

Checking a citation by hand does not scale past a handful of documents. For how a structured review process handles it at volume, see what a risk-ranked AI findings report should actually contain.