Skip to content
Rush Commerce
AI & Automation3 min read

Anthropic's automated researchers: write the benchmark first

Anthropic's automated alignment researchers beat 28 human experts across 10 benchmarks at $4/hour. The lesson for your business: automation needs a scoreboard.

Anthropic published a paper on August 28 showing that Claude can run the alignment research loop on its own — read the literature, propose a training method, train the model, score it, keep what worked, discard what didn't. It beat a panel of experienced human researchers. The interesting part for an operator is not the self-improvement angle everyone is chasing. It's the precondition: the automated researcher only works where somebody already built a benchmark. That constraint applies to every automation you are thinking about buying.

What actually happened

The Anthropic Alignment Science team pointed automated alignment researchers (AARs) at 10 categories of model failure — sycophancy, jailbreaks, prompt injection, power seeking, deception, hallucination, social bias, privacy violation, reward hacking, and concealing uncertainty. Each failure got five AARs running in parallel. Each iteration trains for roughly 30 minutes on a single H200, starts fresh, and carries continuity through persistent memory files. The loop runs up to 48 hours or until the score plateaus.

On all 10 failures, the top method on the leaderboard beat the untrained baseline on a held-out benchmark. The discovered methods kept working on models up to 4.7× larger than the target. Against humans, the best AAR methods outperformed ideas submitted by 28 experienced researchers, and passed the human proposals after about six hours of hill-climbing. TechCrunch reports the AAR runs at roughly $4 per hour against $150 per hour for a human researcher.

Anthropic is explicit about the limit, and it is the whole story: the system is only as good as the benchmark. If the benchmark does not measure what you actually care about, the loop optimizes a number and you get nothing.

Why the benchmark matters for your business

Every automation pitch you hear this year has the same hidden dependency. An agent that drafts your quotes, triages your inbox, or reconciles your invoices can improve on its own only if something can tell it whether the last attempt was better than the one before. Most small businesses have no such thing. The quality bar lives in one person's head, and that person reviews output by feel.

That is the actual bottleneck, and it is not a model problem. Anthropic spent the expensive part of this work building measurable failure categories. You have to do the same thing at your scale before an agent can get better at anything.

Start smaller than you think. Take the workflow you most want automated and write down 20 real cases from the last quarter with the answer you would have accepted. That file is worth more than a vendor demo. It tells you today whether a tool is good enough to deploy, it tells you next quarter whether a model upgrade helped or hurt, and it is the thing that lets you run an improvement loop instead of a vibe check.

The $4-an-hour researcher is not the headline. The headline is that it needed a scoreboard, and so do you.

Key takeaways

  • Anthropic's automated alignment researchers improved performance on all 10 targeted failure categories against a held-out benchmark, and the methods held on models up to 4.7× larger
  • The best automated methods beat proposals from 28 experienced human researchers after roughly six hours of iteration
  • Anthropic's stated limit: the loop is only as good as the benchmark it optimizes against
  • Before automating a workflow, write 20 real cases with accepted answers — that file is your scoreboard and your upgrade regression test

Can you prove your AI tool got better last month, or do you just feel like it did? We build the eval sets and measurement harnesses that turn "the AI seems fine" into a number you can act on — before you spend anything on the automation itself. See how we work.

Sources: Anthropic Alignment Science, TechCrunch.

  • #ai-agents
  • #benchmarks
  • #anthropic
  • #automation
  • #evals
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.