Anthropic's bio classifiers were off for 11 months
Anthropic's August 2026 risk report says its blocking biological classifiers were disabled on human-feedback vendor traffic for 11 months. Log your guardrails.
Anthropic published its second company-wide risk report on August 14, and the headline most people picked up was a rating change: catastrophic misalignment in high-stakes settings moved from "very low" to "low." That is not the interesting part. The interesting part is buried in the incident disclosures — Anthropic's blocking biological classifiers were switched off across an entire traffic path for eleven months, and the same flag that disabled the filter also disabled the logging.
What actually happened
The report covers February 24 through July 15, 2026. In it, Anthropic discloses that human-feedback vendor traffic — the contractor pipeline that generates preference data for training — ran without its blocking biological classifiers from May 2025 through April 2026. Per The Next Web's read of the report, that covered roughly 133 million exchanges with about 50,000 contractors working through external vendors, many of whom did not screen their workforce to the standard Anthropic's own threat model assumes.
The mechanism matters more than the number. A single disabled flag turned off filtering and logging. Flagged traffic was not recorded and never reached any review mechanism. So for eleven months there was no signal that the control was missing — not a degraded signal, not a noisy one, none.
The retroactive review used Claude Sonnet 5 to scan the affected conversations. It flagged 1,197 transcripts as high-risk; 757 came from Anthropic's internal teams, and all but 62 of those traced to deliberate red-teaming. Staff reviewed the remaining 62 and found no clearly problematic misuse, though a handful were potentially dual-use. No customers were affected and the gap is remediated. The risk report also discloses an unreleased internal model — "Model 2" — more capable than Anthropic's shipping frontier model, with no release plans.
Why this matters for your AI governance
Anthropic caught this, wrote it down, and published it. That is the behavior you want from a vendor. Copy the lesson, not the anxiety.
A guardrail you cannot observe is not a guardrail. If your DLP rule, your prompt filter, or your approval gate fails closed silently, you will find out from an auditor. Every control you deploy needs a counter that increments — blocked requests, flagged sessions, denied tool calls. A number that sits at zero for a month is either a clean month or a broken control, and you cannot tell which without a heartbeat.
The bypass is always the side channel. The main product path was fine. The failure was in an internal pipeline run by contractors through third-party vendors. Your equivalent: the batch job, the staging key that still has prod scope, the agent your ops team wired up outside the reviewed stack. Inventory the paths that reach your model without going through your front door.
Ask your vendors for this document. Anthropic's is public. Most aren't. "Show me your last incident review" is a fair procurement question, and the answer tells you more than a SOC 2 logo does.
Key takeaways
- Anthropic's August 14 risk report covers February 24 – July 15, 2026 and raises catastrophic misalignment risk from "very low" to "low" on uncertainty, not a failed test
- Blocking biological classifiers were disabled on human-feedback vendor traffic from May 2025 to April 2026 — roughly 133 million exchanges with about 50,000 contractors
- The same disabled flag killed both filtering and logging, so nothing was recorded or escalated for eleven months
- Retro review flagged 1,197 high-risk transcripts; most traced to internal red-teaming, and staff found no clearly problematic misuse in the remainder
- Give every guardrail a counter you can watch — a silent control and a missing control look identical
- Audit the paths that reach your models outside the main product flow: batch jobs, contractor tooling, internal agents
Most AI controls we inherit have no telemetry at all. We wire filters, approval gates and audit logs so you can prove what your systems blocked — and notice the day they stop blocking anything. See how we instrument AI systems or book a governance review.
Sources: Anthropic Risk Report: August 2026, The Next Web.
- #anthropic
- #ai-governance
- #guardrails
- #logging
- #vendor-risk
Tommy Rush — Founder, Rush Commerce
Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More
Get The Rush Report weekly — one email, zero fluff.
Keep reading
Qwen3.8-27B: the open model you can actually self-host
Alibaba shipped Qwen3.8-27B under Apache 2.0 with a 262K context window. Unlike the 2.4T Max, this one fits on hardware you can rent — here's what that buys you.
Read itNvidia's 13F: $30B in Intel, $21B in SpaceX
Nvidia's Q2 13F shows $30B in Intel and $21B in SpaceX — a chip vendor holding equity in both a supplier and a customer. What that means for your AI costs.
Read it