Skip to content
Rush Commerce
AI & Automation3 min read

Grok 4.6: same Elo, half the steps — price the agent loop

SpaceXAI shipped Grok 4.6 at $2/$6 per million tokens. The number that matters isn't the price — it's how many steps your agent takes to finish the job.

SpaceXAI released Grok 4.6 yesterday, built explicitly for long-running agents. Pricing held at $2 per million input tokens and $6 per million output tokens. Every writeup is leading with that price. The more useful number is buried one layer down: on a multi-step agent benchmark, Grok 4.6 finished tasks in roughly half the steps of the model it tied with. If you run agents, step count is your bill — not the sticker price per token.

What actually happened

Per SpaceXAI's announcement, Grok 4.6 shipped August 12 into Cursor and Grok Build, plus the API console and partners including OpenRouter, Vercel, and Cloudflare. Pricing stays at $2/$6 per million tokens, with a faster variant at double that.

The independent picture, reported by The Decoder from Artificial Analysis data:

  • Intelligence Index: 61, a five-point jump over Grok 4.5 and a tie with GPT-5.6 Sol. Claude Opus 5 sits at 63, Fable 5 at 62.
  • GDPval-AA v2: Grok 4.6 and Claude Opus 5 both scored 1,753 Elo. Opus 5 got there in about 103 steps. Grok 4.6 took about 53.
  • Competitor list pricing runs $5/$25 for Opus 5 and $5/$30 for GPT-5.6 Sol.

SpaceXAI's own scorecard adds CursorBench 3.2 at 69.9%, DeepSWE 1.1 at 65.9%, and FrontierCode 1.1 at 61.3%. Those are vendor-run. Treat them as marketing until someone outside the company reproduces them.

Why agent economics matter more than token price

A chat request is one call. An agent is a loop — read, plan, call a tool, read the result, try again. You pay for every turn, and every turn re-reads the context that came before it. Two models at the same quality bar can differ 2x in how many turns they need, and that difference compounds against a context window that keeps growing.

So the comparison that actually predicts your invoice is cost per completed task, not cost per million tokens. A model at half the sticker price that needs twice the steps is a wash — and worse, it's twice the latency and twice the surface area for something to go sideways mid-run.

One more thing worth pricing in: the 2x included usage in Grok Build and Cursor is a first-week promotion. Whatever your team measures during that window is not what week three costs. We've written before about token caps turning into a meter you didn't budget for — same trap, different wrapper.

Key takeaways

  • Grok 4.6 launched August 12 at $2/$6 per million tokens; the fast variant is double
  • It scored 61 on the Artificial Analysis Intelligence Index, tying GPT-5.6 Sol; Opus 5 leads at 63
  • On GDPval-AA v2 it matched Opus 5's 1,753 Elo in ~53 steps versus ~103
  • Benchmark your agents on cost per completed task, not price per million tokens
  • The 2x usage bonus in Cursor and Grok Build lasts one week — measure your baseline after it ends

Don't know what your agents actually cost per job? We instrument agent workflows so you can see cost, step count, and failure rate per task — and swap the model underneath when the math changes. See how we build AI automation, or run the numbers on your own process first.

Sources: SpaceXAI — Introducing Grok 4.6, The Decoder.

  • #grok
  • #ai-agents
  • #model-pricing
  • #agent-economics
  • #coding-agents
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.