Skip to content
Rush Commerce
Software & Dev3 min read

AI guardrails block security work: keep a second path

Offensive security researchers say OpenAI and Anthropic guardrails now block legitimate defensive work. Your vendor's safety policy is a dependency — plan for it.

The people you hire to break your systems before someone else does are increasingly getting refused by the AI tools they use to do it. In reporting published July 23, TechCrunch spoke to a half-dozen offensive security researchers who say AI guardrails at OpenAI and Anthropic are now blocking legitimate vulnerability research — the same work that finds the bug in your stack before an attacker does. If you buy security services, this is a supply-chain issue wearing a safety label.

What actually happened

TechCrunch interviewed researchers including Mark Dowd, Chris Anley of NCC Group, Chris Thompson of RemoteThreat, Paolo Stagno of Crowdfense, and Giuseppe Cali, plus an anonymous researcher at a smartphone manufacturer. The common complaint: models refuse or degrade on tasks that are indistinguishable from defensive work.

The structural problem is the one Anley names directly — the same tool is both offensive and defensive, and "the two can't really be unpicked." Fuzzing a parser, reversing a binary, writing a proof-of-concept to prove a bug is real: none of these have a safe version and a dangerous version. They have one version, used by two kinds of people.

Vendors have responded with vetted-access programs, which several researchers find unsatisfying — Stagno described them as treating customers like children who need babysitting. Dowd's objection is about who's deciding: he said he isn't comfortable with large companies "making arbitrary decisions about what is safe in security." Thompson went further, arguing the guardrails do more harm than good.

This lands in a year when model access has already proven conditional. A June 2026 report claimed guardrails on Anthropic's Mythos and Fable models could be bypassed, and Fable has been subject to US export restrictions. The pattern is consistent: the capability you rely on is governed by a policy you don't set.

Why it matters for your business

You probably don't run a vuln research team. You almost certainly buy from someone who does — your pen test vendor, your managed security provider, the researcher who filed your last disclosure. Their tooling just got less reliable, and the cost shows up in your engagement as slower turnaround or thinner coverage, not as a line item that says "the model refused."

Two practical moves. First, ask the question in procurement: when your security vendor's AI tooling refuses a task, what's their fallback? If the honest answer is "we wait for the appeal," that's a schedule risk you now own. Vendors running open-weight models locally for the parts that get refused have a real answer; vendors with a single frontier API do not.

Second, apply the same test inward. Every AI-dependent workflow you run — code review, log triage, patch verification — has a refusal mode you haven't hit yet, because the policy that would trigger it hasn't shipped yet. The question isn't whether your provider's rules are right. It's whether your process stops when they change. Keep a second path: a local model, a deterministic tool, or a documented manual fallback for anything that touches a deadline.

Key takeaways

  • Offensive security researchers told TechCrunch (July 23) that OpenAI and Anthropic guardrails are blocking legitimate vulnerability research
  • The core problem is structural: fuzzing, reversing, and proof-of-concept work have no separable "defensive-only" version
  • Vendor vetted-access programs exist but researchers object to the gatekeeping and the latency it adds
  • Ask your security vendor what their fallback is when a model refuses — and build a second path for every AI-dependent workflow you run on a deadline

What happens to your workflow the day the model says no? We build AI systems with the refusal path designed in — local models, deterministic tools, and documented manual fallbacks — so a policy change upstream doesn't stop your week. See how we build systems you own or tell us which workflow can't afford to stall.

Sources: TechCrunch — How AI guardrails are impeding the work of offensive cybersecurity researchers.

  • #ai-guardrails
  • #security
  • #vendor-risk
  • #penetration-testing
  • #portability
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.