Skip to content
Rush Commerce
Tools & Teardowns3 min read

Pangram raises $9M: AI detection claims need testing

Pangram raised $9M from Menlo Ventures and shipped Pangram 4 plus an image detector. Every accuracy number is vendor-supplied. Here's how to test it on your own data.

Pangram announced a $9 million round on July 29 alongside Pangram 4, its new text detector, and a research preview of an image detector. Nine days earlier, Epoch AI published testing showing the previous Pangram model missing a quarter of style-imitated AI text. Both things are true, and the gap between them is the whole story for anyone thinking about using AI detection as a control.

What actually happened

Per TechCrunch and SiliconANGLE, the round was led by Menlo Ventures with Haystack, ScOp, Script Capital, and Cadenza participating. Pangram sells a $20/month web subscription, a Chrome extension that labels posts inline on X, LinkedIn, Substack, Reddit, and Medium, and an enterprise API. Named users include Substack, which has it integrated, and Quora.

The product claims: Pangram 4 detects AI-assisted and mixed human-AI writing at over 99% accuracy, with a false-positive rate of 0.0041% — roughly one wrongly-flagged document in 24,000. TechCrunch describes the same figure more conservatively as about one in 10,000. The company also says v4 is more robust against "humanizer" tools. Pangram Image, in research preview, works on pixel-level statistical distributions and is said to catch output from GPT Image, Gemini Nano Banana, Midjourney, FLUX, and Grok Imagine, plus video from Kling, Seedance, Veo, and Wan.

Every one of those numbers is internal. There is no third-party evaluation of Pangram 4 yet. There is a third-party evaluation of Pangram 3: Epoch AI found that when a model was handed a writing sample to imitate, Pangram v3.3.2's miss rate went from 0.7% to 10% — and to 25% on scientific writing.

Why detector accuracy matters for your business

A funding round is not a validation. If v4 genuinely closes the style-imitation gap, that's real progress and we'll say so when someone independent measures it. Until then, the reasonable posture is unchanged: a detector output is a signal, not a verdict.

The false-positive rate is the number that should decide your policy, and it only means something against your distribution. 0.0041% measured on a benchmark corpus is not 0.0041% on the writing of a bilingual employee, a technical writer with a clipped house style, or a contractor who drafts in bullet points. Test it yourself before you wire it to a consequence:

  • Assemble a holdout set from your own archives — 200 documents you know are human, written before mid-2022 if you want to be strict, and 200 AI drafts produced in your house style with your actual prompts.
  • Run both through the tool and compute your own FPR and FNR. Two hours of work, and it will not match the marketing page.
  • Decide the consequence before you see the score. "Flags route to a named human reviewer" is a policy. "Flagged work is rejected" is a lawsuit waiting for its first false positive.

And keep building the thing detectors can't replace: provenance at the source. Draft history, commits, revision trails, disclosure requirements in the contract. A document with no edit history tells you more than any classifier will.

Key takeaways

  • Pangram raised $9M led by Menlo Ventures on July 29, shipping Pangram 4 and a research-preview image detector
  • Claimed 0.0041% false-positive rate (~1 in 24,000); TechCrunch characterizes it as ~1 in 10,000 — all vendor-supplied
  • Independent testing of the prior model found 10% miss rates on style-imitated text, 25% on scientific writing
  • Image model covers GPT Image, Gemini Nano Banana, Midjourney, FLUX, Grok Imagine, and several video generators
  • Operator move: build a 200/200 holdout from your own archives, measure your own error rates, and route flags to a human — never to an automatic rejection

Governance that survives a false positive. We build review workflows around provenance and named approvers, not around a confidence score you can't audit. See how we build it.

Sources: TechCrunch, SiliconANGLE, Epoch AI.

  • #ai-detection
  • #funding
  • #content-provenance
  • #governance
  • #verification
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.