Skip to content
Rush Commerce
Tools & Teardowns3 min read

AI text detectors miss up to 29% of styled AI writing

Epoch AI tested Pangram, GPTZero, and Originality.ai. Give a model a writing sample to imitate and detection collapses. Why you can't govern with AI detectors.

If your policy for AI-generated content is "we run it through a detector," you don't have a policy. New testing from Epoch AI, covered July 19 by The Decoder, shows that AI text detectors fall apart the moment a model is given a writing sample to imitate — which is exactly how anyone actually uses them.

What actually happened

Epoch tested three of the most widely deployed detectors: Pangram v3.3.2, GPTZero (model 2026-05-11-base), and Originality.ai Turbo 3.0.2. Against plain, unprompted AI output, they were near-flawless — Pangram missed 0.7% of it. That's the number vendors quote.

Then Epoch handed the model a reference sample and asked it to write in that author's style. Miss rates jumped to 10% for Pangram, 11% for GPTZero, and 18% for Originality.ai — about 13% on average across the three.

On scientific writing, the genre where these tools probably see the heaviest real-world use, it got worse: 25% missed by Pangram, 24% by GPTZero, and 29% by Originality.ai. Roughly a quarter to a third of style-imitated AI text sailed straight through.

The errors run the other way too. Originality.ai flagged 19 of 495 genuine human passages as AI — a 3.8% false-positive rate. At any real volume, that's people being accused of something they didn't do.

Why detector accuracy matters for your business

The failure mode isn't random. It's adversarial. A detector that catches lazy output and misses careful output is a filter on effort, not on provenance — and anyone with a motive to evade it clears the bar with one extra sentence in the prompt. Building a hiring screen, a contractor QA gate, or an academic integrity process on that is worse than having nothing, because it manufactures false confidence in both directions: undetected AI that passes, and real humans flagged as frauds.

The fix is boring and it works: verify provenance at the source, not the output. Require drafts, revision history, and version control from contractors — a document with no edit history is a signal a detector will never give you. Judge work against outcomes you can measure: does the code pass the tests, does the copy convert, does the analysis hold up when someone checks the numbers. Put a named human reviewer on anything that ships. And if you need an AI-use policy, write one about disclosure and accountability rather than detection, because disclosure is enforceable and detection isn't.

We build automation with AI in it every day. That's exactly why we don't trust a classifier to tell us who wrote something. Verification belongs in the workflow, not bolted onto the end.

Key takeaways

  • Epoch AI tested Pangram v3.3.2, GPTZero, and Originality.ai Turbo 3.0.2 against plain and style-imitated AI text
  • On plain AI output detection was near-flawless; with a style sample, miss rates hit 10–18%
  • On scientific writing, 24–29% of style-imitated AI text went undetected
  • Originality.ai false-flagged 3.8% of genuine human passages (19 of 495) as AI
  • The operator move: verify provenance via drafts and version history, judge by measurable outcomes, and write policy around disclosure — not detection

Need a review process that actually holds? We build workflows where provenance, approvals, and outcome checks are part of the system — not a classifier you hope is right. See how we build it.

Sources: Epoch AI — AI text detection challenges, The Decoder.

  • #ai-detection
  • #content-ops
  • #governance
  • #verification
  • #tooling
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.