Treadstone Associates
Article · 8 min read

Why AI detectors get it wrong

An AI-detection tool works by pattern-matching against how known models tend to generate content — which means it is structurally always chasing yesterday’s model, and it produces two different kinds of wrong answer for two different reasons.

Treadstone Associates · Updated 2026

Key takeaways

  • • Watermark-based detection only works if the specific tool embedded a watermark in the first place — it produces a false “not AI” reading on anything from a different source.
  • • Pattern-based detectors face the opposite problem: they key off stylistic regularities that also occur naturally, so ordinary writing can be flagged as AI-generated.
  • • Content Credentials’ own framing is that provenance tools exist for good actors to prove authenticity, not to catch every bad actor’s fake — a structural ceiling the coalition names itself.
  • • No accuracy rate for any AI detector — Canadian or otherwise — is published by a source found here; a detector’s own marketing is not a substitute for one.

Two different failure directions

A false negative happens when content is AI-generated but a tool reports that it isn’t — the routine case whenever content comes from a model the detector was never built to recognize, or when a genuine watermark has been stripped by ordinary processing, covered in the companion piece on metadata loss. A false positive runs the other way: content that is not AI-generated gets flagged anyway. Pattern-based approaches key off stylistic regularities that also occur naturally in human writing — short declarative sentences, common transition words, a narrow vocabulary — which are not unique to machine output.

Why watermark detection specifically is narrow

Google describes SynthID’s coverage precisely: the watermarks are “embedded across Google’s generative AI consumer products.” (deepmind.google) That means a Gemini-based check can only ever confirm Google’s own watermark — a “no” result answers the question “was this watermarked by Google,” not the broader question “was this AI-generated at all.” Content from any other model returns no signal either way, which a detector’s summary output can easily present as a false “clean” result if the person reading it doesn’t know the distinction.

Why provenance-based tools face a different limit

Content Credentials names its own ceiling directly: “While there will continue to be bad actors who seek to label synthetic content as authentic, our goal is to provide good actors with a way to demonstrate the authenticity of their content.” (contentcredentials.org) That is a design choice stated plainly, not an oversight: the system assumes good-faith participation from whoever creates and handles the file, and it is not built to unmask someone actively trying to defeat it by stripping or forging a record.

Why the arms race structurally favours whoever generates content

A pattern-based detector has to be trained or tuned against models that already exist by the time it ships. A generation model released after that point owes it nothing — there is no mechanism by which a newer model is bound to keep producing the patterns an older detector learned to recognize, and every incentive for a model built for general writing or image tasks to drift stylistically as it improves for unrelated reasons. That asymmetry is structural, not a temporary gap that better engineering closes: the detector is always describing the last generation of models it was shown, and the field it is describing keeps moving while it is being built.

Even Canada’s own policy treats detection as unfinished

The federal Voluntary Code of Conduct’s own wording gives the game away: it asks a manager to build “a reliable and freely available method to detect content generated by the system, with a near-term focus on audio-visual content (e.g., watermarking)” — near-term, not permanent, and audio-visual first because text detection is the harder, later problem. (ised-isde.canada.ca) A federal policy document conceding that detection is presently only a near-term priority is a reasonable signal that no one, government included, currently claims a solved, general-purpose detector exists.

What this means practically

Treat a detector’s output as one weak signal among several — source reputation, corroboration, a provenance record if one is present — rather than a verdict on its own. Canada’s Centre for Cyber Security gives the same underlying advice for AI content generally: outputs “can be incorrect,” “might not make sense,” and “can be biased,” so a reader should “always be aware of and validate your sources to verify whether the content being presented is accurate.” (cyber.gc.ca) The companion piece on provenance versus detection sets out why recording origin at creation is structurally more reliable than guessing it afterward; a detector’s report is exactly the “guessing afterward” half of that comparison.

A worked example

A hiring team runs a candidate’s written response through a free online “AI detector” that reports 92% AI-generated. No Canadian or vendor source found here documents that specific tool’s actual accuracy rate, and the stylistic markers that trip pattern-based detectors — short sentences, common transition words — also occur in writing by non-native English speakers and in text a human wrote and then ran through an AI tool only for grammar. The reported percentage is not itself evidence of anything actionable; treating it as a verdict is exactly the failure mode this article describes. A defensible next step is to ask the candidate to talk through their own answer rather than act on the score directly — a conversation tests understanding in a way no detector output does, and it does not risk penalizing a candidate for a false positive the tool cannot itself rule out.

If that hiring team is the one using the AI detector rather than the one being screened, Ontario law has begun to reach the disclosure side of this scenario, if not the detector itself. Effective January 1, 2026, the Employment Standards Act, 2000 requires that “every employer who advertises a publicly advertised job posting and who uses artificial intelligence to screen, assess or select applicants for the position shall include in the posting a statement disclosing the use of the artificial intelligence.” (ontario.ca, ESA 2000 s.8.4) The duty is only to disclose that AI is being used to screen at all — it says nothing about which detector, or whether its output was accurate, which is exactly the gap this article describes.

Related: combining a detector check with other signals, why recording origin beats testing the artifact afterward, the mechanism a watermark-based detector actually checks

Common questions

Are AI detectors ever reliable?

For a narrow claim — does this carry a specific vendor’s watermark — yes, within that vendor’s own products. For the broader claim of whether something was written or made by any AI at all, no accuracy figure from a Canadian or vendor source fetched here supports treating a detector’s output as reliable on its own.

Why would a detector flag something a human actually wrote?

Because pattern-based detectors key off stylistic regularities — sentence length, common phrasing, structure — that occur naturally in human writing too, especially concise or formulaic writing. The pattern is a correlation, not a fingerprint unique to AI.

Should a business rely on a detector’s score to make a decision about someone?

Not on its own. Because neither watermark-based nor pattern-based detection is documented here as reliable in isolation, a detector’s score is better treated as one input to review, not the basis for a decision about a person or a piece of content by itself.

Will detectors eventually get accurate enough to trust on their own?

Nothing sourced here supports either a yes or a no. The structural asymmetry — a detector trained against models that already exist, facing a field that keeps producing new ones — does not obviously resolve with more engineering effort, which is a reason to design a review process around corroboration now rather than wait for a detector good enough to replace it.

Where a detector fits into an operational review process matters more than the detector itself.

Building a review process that doesn’t over-trust a single detector’s score is an operations problem, not a technology purchase.