Treadstone Associates
Guide

How to sanity-check an answer from AI

Nothing in a typical AI chat interface tells you how confident the system actually is — the tone is the same whether it is right or badly wrong. This guide is the sequence of checks that separates an answer you can rely on from one you cannot, in the order that catches the most problems for the least effort.

Treadstone Associates · Updated 2026

Key takeaways

  • • Fluency is not a confidence signal. A wrong answer and a right one are written in exactly the same tone, because the model is optimising for plausible language, not flagging its own uncertainty.
  • • Canada’s Cyber Centre states plainly that generative AI output “can be incorrect”, “might not make sense”, and “can be biased” — and that you should “always be aware of and validate your sources.”
  • • Checking a citation the AI gives you means opening the source and reading it — not asking the AI whether it is sure, which tests nothing independent.
  • • How much checking a specific answer needs should scale with what happens if it is wrong, not be the same routine every time.

STEP 01 OF 11

Separate “this sounds right” from “this is right”

A generative AI system is built to produce plausible, well-formed language — that is a separate property from producing correct language, and the two can diverge without any change in tone. Canada’s Cyber Centre defines generative AI as a system that “generates new content by modelling features from large datasets”, and states directly that its output “can be incorrect”, “might not make sense”, and “might not take certain factors into account” — see ITSAP.00.041.

The practical habit this creates: read a confident answer the same way you would read a confident answer from a stranger with no track record on the topic — useful as a starting point, not as a conclusion.

STEP 02 OF 11

Do not treat the AI’s own confidence language as a signal

Phrases like “I am confident that” or a caveated “this may vary” are generated the same way as the rest of the answer — they reflect patterns in how confident-sounding or hedged text is usually phrased, not a genuine internal measurement of how likely the answer is to be correct. Asking the system “are you sure” and treating a reaffirming answer as verification tests nothing independent of the first answer.

If a tool does expose an actual calibrated confidence score, that is different information and worth using — the point is not to distrust every signal, it is to check whether a given signal is actually measuring something before relying on it.

STEP 03 OF 11

Apply the working definition of “valid”: does independent evidence back it up

The U.S. NIST framework — cited by Canada’s privacy regulators, see Step 7 — defines validation as “confirmation, through the provision of objective evidence, that the requirements for a specific intended use or application have been fulfilled” (from AI Risks and Trustworthiness). Translated into a habit: find evidence that exists independently of the AI’s own answer, not another AI answer restating the same claim.

Independent evidence means a primary source, a document you can open, a number on a page you can read yourself — not a second chatbot session, which likely draws on overlapping training data and can fail in a correlated way.

STEP 04 OF 11

Trace any number or fact back to a primary source, yourself

If the answer includes a statistic, a rule, or a specific figure, find where that figure actually lives and read it there. See how to read AI statistics without being fooled for what tends to go wrong specifically with numbers. A figure that cannot be traced to a real, checkable source is not a fact you have — it is a plausible-sounding claim you have not yet verified.

This applies even when the AI names a source. Naming a source and accurately representing what that source says are two different claims, and only one of them is checked by seeing a citation appear on the screen.

STEP 05 OF 11

Open the cited source and read the actual sentence

This is the single highest-value check and the one people skip most often. A citation that resolves to a real, live page is not the same as a citation that says what the AI claims it says — an AI can summarise a source slightly wrong, blend two sources into one claim, or attribute a real sentence to the wrong page entirely.

Open the link. Find the specific sentence. Confirm it says what you were told it says, word for word if you intend to repeat it as a quotation. This single habit catches a meaningful share of AI errors that every other check in this guide misses.

STEP 06 OF 11

Ask the same question a different way and compare

Rephrase the question, or ask it from a different angle, and see whether the substance of the answer holds up. NIST’s definition of reliability is the “ability of an item to perform as required, without failure, for a given time interval, under given conditions” — from the same NIST page. An answer that changes substantively when asked a slightly different way is telling you something real about how much weight it can bear.

See also why AI answers the same question differently — some variation is normal and does not itself indicate an error, but a substantive contradiction between two phrasings of the same question is a genuine warning sign.

STEP 07 OF 11

Watch for the trade-off between confident and cautious

The same NIST source notes that “accuracy and robustness… can be in tension with one another” in AI systems. In practice, a system tuned to avoid confidently wrong answers will sometimes hedge on things it actually could answer well, and a system tuned to be maximally helpful will sometimes answer confidently past the edge of what it actually knows.

Neither failure mode announces itself. A vague, hedged answer is not automatically the safer one — it can simply be unhelpful. Judge the specific answer in front of you rather than assuming a cautious-sounding tool is a reliably correct one.

STEP 08 OF 11

Know where the Canadian chain for any of this actually comes from

This is not a US framework operating on its own authority in Canada. Canada’s joint federal-provincial-territorial privacy principles for generative AI cite the NIST framework directly: footnote 16 of the OPC’s generative AI principles (2025-05-06) reads, “For more information on validity and reliability in AI systems, see the NIST AI Risk Management Framework”, attached to the principle that organisations “evaluate the validity and reliability of the generative AI tool for the intended purpose.”

That is the honest chain: a Canadian regulator principle, pointing to a US technical framework for the detail. Cite both together, and label NIST as American, if you are writing this up for someone else.

STEP 09 OF 11

Scale the checking to what happens if you are wrong

Not every answer deserves the full sequence above. An AI-drafted summary you will personally review before sending needs less scrutiny than a number that will go into a customer-facing document unread. Work out what a wrong answer would actually cost in this specific case, then match the depth of checking to that — see how to decide what a human must approve for the same logic applied to an automated process rather than a single answer.

Checking everything to the same exhaustive standard is not realistic day to day, and pretending otherwise usually means nothing gets checked properly at all.

STEP 10 OF 11

Treat the answer as a lead, not a conclusion, when the stakes are real

For anything that will be repeated to a client, filed, or acted on with real consequences, the AI answer’s real job is to point you toward where to look, not to be the final word itself. Use it to identify the likely source, the right search term, or the shape of the answer — then confirm the substance the way you would confirm anything else you intended to rely on.

This is not a criticism of the tool. It is simply an accurate description of what a fluent, well-sourced-sounding answer is and is not evidence of.

STEP 11 OF 11

Notice when a question is outside what the system could plausibly know

Some questions are unanswerable by a given tool in principle, not just unreliably answered — anything genuinely recent, anything specific to your own private records the tool has no access to, anything that depends on a change in the law made after the system’s training ended. A confident-sounding answer to one of these questions is a stronger warning sign than a confident answer to a question the system plausibly could know.

Before trusting an answer, ask yourself where the system could actually have learned this. If the honest answer is “nowhere obvious,” treat the response as a guess dressed as an answer, however specific it sounds.

Common mistakes

Treating a confident tone as a confidence score. The two are unrelated. A wrong answer reads exactly as assured as a right one, because both are generated the same way.

Asking a second AI to check the first AI's answer. Different tools, and even different sessions with the same tool, can share correlated blind spots drawn from overlapping training data. This is not the same as an independent check.

Seeing a citation and stopping there. A link that resolves and a link that actually supports the claim next to it are different things. Open it and read the sentence.

Asking a leading question and treating the answer as confirmation. A question phrased to expect a particular answer tends to get one. Rephrase neutrally, as in Step 6, rather than testing whether the system will agree with what you already believe.

Trusting a specific-sounding answer to a question the system could not plausibly know. Specificity is not evidence of access to real information — a system can produce a precise-looking number or date for a question it has no way of actually knowing the answer to, as in Step 10.

A short checklist, in order of effort

  • • Cheapest: rephrase the question and see if the substance holds (Step 6).
  • • Cheap: ask whether the system could plausibly know this at all (Step 10).
  • • Cheap: open any cited source and read the actual sentence (Step 5).
  • • Moderate: trace any number to a primary source you can read yourself (Step 4).
  • • Only when the stakes justify it: treat the whole answer as a lead and independently reconstruct the conclusion (Step 9).

Running the cheap checks by default and reserving the expensive ones for high-stakes answers is more sustainable than trying to run every check on everything — and it is closer to what most people can actually keep up in practice. None of these checks require special tools or technical knowledge; they require deciding, in advance, that a fluent answer is not the same thing as a verified one.

Frequently asked

Can I trust a citation an AI gives me?

Only after you open it and confirm the specific sentence it is attributed to actually appears on that page. A citation that resolves to a live, real page is a weaker guarantee than it looks — the AI can still misattribute or slightly misquote what that page says.

Does asking the AI “are you sure” actually help?

Not on its own. The follow-up answer is generated the same way as the first one and is not an independent check — it can reaffirm a wrong answer just as confidently as it stated it the first time.

Is a longer, more detailed answer more likely to be correct?

No reliable relationship exists between length and accuracy. Detail can make an answer feel more trustworthy without adding anything that was actually verified — treat length as a writing style, not evidence.

What is the single fastest check if I only have time for one?

Rephrase the question and see whether the substance holds up, per Step 6. It costs almost nothing, and a substantive contradiction between two phrasings of the same question is one of the more reliable warning signs available without leaving the conversation.

Sanity-checking one answer is different from keeping a whole workflow correct over time.

Once AI output feeds a running process, the checks need to run continuously, not once.