Skip to content
Rush Commerce
Tools & Teardowns3 min read

Faraday: a 27B agent beat frontier models at replication

Inherent's 27B Faraday agent beat Claude Opus 4.8 and GPT-5.5 at replicating research papers. What a small-model win means for your AI spend.

Inherent, a London lab founded by Google DeepMind alumni, released Faraday, a 27B-parameter agent trained to reproduce the findings of published research papers. The company reports it outperforms Claude Opus 4.8 and GPT-5.5 at that job. The interesting number is not the win. It is 27 billion — a model small enough to run on hardware you can actually rent, beating systems an order of magnitude larger at a task that is mostly persistence and judgment.

What actually happened

Per Inherent's research write-up, Faraday ships alongside Replica, a benchmark of 310 tasks drawn from 100 machine-learning and AI-for-science papers spanning natural language processing, materials science, and weather forecasting. Each task asks an agent to reproduce a figure from a paper without access to the original plot, under a limited time and compute budget. The point is not to recall a known answer — it is to recover the experimental work papers leave out.

Training was long-horizon reinforcement learning with coding agents as a tool. The reward signal is the hard part, and Inherent's answer is an auto-generated rubric-based judge, multi-sample aggregation, and turn-level credit assignment rather than leaning on a single LLM judge. The arXiv paper runs 47 pages and 12 figures.

Inherent says Faraday produces more faithful replications than both frontier baselines across every category in Replica, with the widest margins in meta-learning, structural biology, and materials science. TechCrunch reports the base model is Qwen 3.6 at 27B, and that Inherent came out of stealth weeks ago on a $50M seed with roughly a dozen people.

Two caveats worth stating plainly. This is a vendor-built benchmark scored on vendor-defined tasks, and the abstract publishes no headline number — only a directional claim of surpassing the baselines on held-out tasks. Treat it as a signal, not a scoreboard.

Why a 27B agent matters for your business

Task-shaped training beats parameter count on narrow work. Nobody is claiming Faraday is smarter than Opus. They are claiming it is better at one long-horizon task, because it was trained on that task with a reward that measured it. Most of your AI work is narrow too.

The reward signal is the asset, not the weights. Inherent's real product is the rubric-based judge and the 310-task space. If you want an agent that reliably does your intake, your reconciliation, your quoting — the leverage is in defining what "done correctly" means precisely enough to score, not in picking a bigger model.

27B changes where the model can live. A 27B open-weights model runs on a single rented GPU. That moves your inference from a per-token line item on someone else's invoice to a fixed cost you control, and it removes an entire class of vendor-availability risk.

Build your own Replica. The transferable move here is not "use Faraday." It is: assemble 50–300 real tasks from your own operation, write the rubric, and score every model you consider against it. Vendor benchmarks tell you what the vendor optimized.

Key takeaways

  • Faraday is a 27B-parameter agent trained via long-horizon RL to replicate research papers
  • Replica benchmark: 310 tasks from 100 papers across NLP, materials science, and weather forecasting
  • Inherent reports it beats Claude Opus 4.8 and GPT-5.5 across every category — vendor-run, no published headline score
  • Base model is Qwen 3.6 (27B) per TechCrunch; Inherent raised a $50M seed and has roughly 12 people
  • The reusable idea: a scored rubric over real tasks is what makes a small model competitive

The model is the cheap part. The eval is the work. We build agent systems where your tasks are scored against a rubric you own, so swapping a frontier model for a 27B one is a config change and a measurement — not a rewrite. See how we build agent systems, or bring us the workflow you want scored.

Sources: Inherent Labs research, arXiv:2608.13331, TechCrunch.

  • #ai-agents
  • #open-weights
  • #benchmarks
  • #model-cost
  • #evals
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.