Skip to content
Rush Commerce
Tools & Teardowns3 min read

OpenAI edited Astra's benchmarks after launch. Snapshot the numbers.

Fortune tracked GPT-6 Astra benchmark scores changing on OpenAI's live blog post after launch. Vendor pages are mutable — archive the numbers you decide on.

Fortune took timestamped snapshots of OpenAI's GPT-6 Astra announcement blog post and watched the benchmark numbers move. Astra's reported hallucination rate went from 4.2% to 2% and back to 4.2%. Its ARC-AGI-3 score was 98.6% in the embargoed press draft and 99.99% on the live page. Anthropic's Fable 5.1 dropped roughly ten points on FrontierMath Tier 4, then partially recovered. Nobody announced any of it. If you picked a model off that page on launch day, you picked it off a document that was still being edited.

What actually happened

Fortune's reporting is worth reading for the timestamps alone. The specific swings:

  • Astra hallucination rate: 4.2% at 2:23 p.m. ET, still 4.2% at 3:11 p.m., 2% at 5:20 p.m., back to 4.2% now.
  • GPT-5.6 Sol hallucination rate: 12.2%, then 9.4%, now 12.2% again.
  • Astra on ARC-AGI-3: 98.6% in the embargoed pre-publication draft, 99.99% on the published page.
  • Fable 5.1 on FrontierMath Tier 4: 87.8%, then 78%, now 83%.
  • Sol on ExploitBench: 5.5%, then 11.5% — OpenAI said it was investigating a reversion after admitting the higher figure reflected a reasoning level not commercially available for that model.

OpenAI's explanation: "Most evaluations have noise within a few percentage points based on the exact checkpoint, scaffold, and evaluation run used in reporting." That is a fair description of eval noise in general. It is a thinner explanation for a competitor's score moving ten points in one direction and a house score halving in the other, on the same page, in the same afternoon.

Why it matters for your business

We have argued before that vendor benchmarks are marketing artifacts, not fitness tests. This adds a mechanical wrinkle: the marketing artifact is a live web page, and it can change under you after you have made a decision based on it.

The practical fix takes about two minutes per decision. When you choose a model — for a coding agent, a support triage flow, a document extractor — archive the source. Save the vendor's page to the Wayback Machine, or print it to PDF with the date on it, and drop it next to the ADR or the ticket where you wrote down why you picked that model. Cite the snapshot, not the URL. When somebody asks in six months why you are paying for Astra instead of the cheaper alternative, you want the numbers you actually saw, not the numbers currently displayed.

Then do the thing the snapshot cannot do for you: run the model on ten to twenty real tasks from your own business, graded on outcomes you can verify. Did the invoice reconcile. Did the test suite pass. Did a human sign off without editing. A vendor page that changes twice in an afternoon is not a source of truth about your workload. It is a source of truth about what the vendor wanted you to believe at 5:20 p.m.

Key takeaways

  • Fortune's timestamped snapshots show GPT-6 Astra's hallucination rate reported as 4.2%, then 2%, then 4.2% again on OpenAI's live announcement post
  • Astra's ARC-AGI-3 figure was 98.6% in the embargoed draft and 99.99% on the published page
  • Anthropic's Fable 5.1 FrontierMath Tier 4 score moved 87.8% → 78% → 83%; OpenAI said it was investigating a reversion on a Sol ExploitBench figure it admitted reflected a non-commercial reasoning level
  • OpenAI attributes the movement to normal evaluation noise from checkpoint, scaffold, and run differences
  • Archive the vendor page you decided on, cite the snapshot in your decision record, and grade candidates on your own tasks

Picking a model on a vendor's number? We build a private eval harness from your real tasks and a decision record that survives the vendor editing its own page. See how we do it or bring us your workload.

Sources: Fortune — OpenAI quietly boosts some of Astra's evaluation metrics.

  • #benchmarks
  • #evals
  • #model-selection
  • #openai
  • #vendor-claims
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.