Claude computed a nine-loop amplitude for ~$100
Anthropic's Claude beat a 2023 physics record using known methods, 96 CPUs and about a week. The lesson for operators: long-horizon agent runs got cheap.
A physicist dared the labs in August. On September 25 Anthropic answered: Claude computed the six-particle scattering amplitude in planar N=4 super Yang-Mills theory at nine loops, one loop past the record SLAC's Lance Dixon and Andy Liu set in 2023. The number that should interest you is not nine. It is $100 — the compute bill for the bootstrap run, 96 CPUs for about a week. A nine-loop amplitude at the price of a mid-tier SaaS seat.
What actually happened
Matt von Hippel, a former theoretical physicist, issued a public challenge on August 7: show that an AI can crack one of the field's open problems using the computing resources an academic actually has. Not a national lab. A grant budget.
Per Anthropic's write-up, researchers Liam Fitzpatrick and Siddharth Mishra-Sharma pointed Claude Fable 5.1, running in the Claude Science harness, at the nine-loop problem. Two approaches were tried. The bootstrap method — Python and SymPy — cost roughly $100 on 96 CPUs over about a week. The indirect form-factor route came in at $1,000 to $2,000. Anthropic's framing is that an end user would have paid one or two thousand dollars all in.
The supervision is the part worth reading twice. The instruction was, effectively, keep working on this until I tell you to stop, give me updates every 4-6 hours. Anthropic describes it as one shot "without any scientific oversight more sophisticated than 'keep going.'"
Dixon, who held the prior record, verified the result. His comment on Claude — "a different kind of transformer model, probably over a million times bigger than our custom one" — is a physicist noting that the machine ran his methods. Anthropic says so plainly: Claude used known methods with more compute than anyone had bothered to throw at them. No new mathematics. Von Hippel's reaction is the honest one: "if you didn't know it could do that because you're still thinking of AI as so error-prone that it's unusable, then this should be your takeaway: it can do this kind of thing reliably now."
Why cheap long-horizon agent runs matter for your business
The unattended week is the product, not the physics. Most businesses still use AI in 30-second increments — a draft, a summary, a snippet. The result here is a model running for a week against a defined objective with a human checking in twice a day. That shape maps directly onto the unglamorous work you have been postponing: reconciling five years of inconsistent SKU data, rebuilding a pricing model from raw invoices, migrating a schema nobody documented.
Your documented method is the asset. Claude did not invent the bootstrap approach. Dixon's group built it over a decade and published it. The model read the papers and pushed the crank one turn further than anyone had funded. If your process lives in one person's head, an agent cannot extend it. Write the method down — that is the input that made this work.
Price the run, not the seat. A hundred dollars of CPU and a week of wall clock is a line item you can approve without a meeting. We keep telling clients to stop budgeting AI per user and start budgeting per job, because the review step is the real cost and the compute usually is not.
Verification is still human. Dixon checked the answer. Nobody shipped this on the model's say-so. Any long-horizon run you start needs a named person who can tell right from plausible, and a check that runs before the output touches a customer.
Don't over-read it. One problem, one field, a well-specified target with a known method and a checkable answer. Your AR aging report is messier than N=4 super Yang-Mills in exactly the way that matters: nobody can verify it in an afternoon.
Key takeaways
- Anthropic published on September 25 that Claude Fable 5.1 computed a nine-loop six-particle amplitude in planar N=4 super Yang-Mills, beating the 2023 eight-loop record
- The bootstrap run cost about $100 — 96 CPUs for roughly a week; the indirect form-factor method cost $1,000–$2,000
- Supervision was "keep going" with updates every 4–6 hours, not step-by-step scientific direction
- Lance Dixon, who held the prior record, verified the result
- Claude used methods Dixon's group developed and published — more compute applied to known math, not new math
- The transferable lesson: unattended multi-day agent runs against a written method are now cheap enough to approve without a business case
- A human still verified the output; build the check step before you start the run
The backlog you keep deferring because it would take someone a week is now a $100 compute job — if the method is written down. We document the process first, then build the agent that runs it unattended, with a review gate you control. See how we build agent systems, or tell us which backlog you'd start with.
- #ai-agents
- #long-horizon-tasks
- #claude
- #automation-roi
- #research
Tommy Rush — Founder, Rush Commerce
Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More
Get The Rush Report weekly — one email, zero fluff.
Keep reading
US and China open an AI incident hotline
The White House says the US and China will run a Super Intelligence Dialogue and a bilateral AI incident channel. What an incident channel implies for your stack.
Read itOpenAI agents posted 53 ChatGPT user images online
OpenAI says its own research agents pushed 53 user-supplied ChatGPT images to public image hosts. Training-data consent is a data-exit path, not a checkbox.
Read it