Skip to content
Rush Commerce
AI & Automation4 min read

Claude Sonnet 5.5 beats Opus 5.5 on coding at half the price

Claude Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0 against Opus 5.5's 66.4%, at $2/$10 per million tokens. Your default model just changed.

Anthropic shipped Claude Sonnet 5.5 today, and the number that matters is not the price — it is that the cheap model now outscores the expensive one on agentic coding. Sonnet 5.5 hits 70.6% on Terminal-Bench 4.0. Opus 5.5, which came out six days ago at twice the token price, hits 66.4%. If you routed your coding agents to Opus last week on the assumption that bigger is better, that assumption is now wrong and expensive.

What actually happened with Claude Sonnet 5.5

Anthropic released Sonnet 5.5 as the second model in the Claude 5.5 family, at $2 per million input tokens and $10 per million output — identical to Sonnet 5's pricing. Cache reads are $0.20 per million. It is available on the Claude Platform, AWS, Google Cloud and Microsoft Azure as claude-sonnet-5-5, with a 1M-token context window and 128K max output per the model docs.

The benchmark jumps from Sonnet 5 are not incremental. Terminal-Bench 4.0 goes from 10.3% to 70.6%. OSWorld 2.1, the computer-use test, from 57.0% to 80.1%. CursorBench 4.0 from 34.1% to 55.5%. GDPval-AA from 1449 to 1844 — Opus 5.5 scores 1846 on the same test. On most evals Sonnet 5.5 now lands within a point or two of Opus 5.5.

Anthropic's cost claim is subtler than a price cut: same sticker price, up to 30% less per task, because the model finishes with fewer tokens and fewer tool calls. Slack measured about 14% fewer output tokens. VentureBeat reports Lovable saw a third fewer tool calls and roughly half as many shell executions, and Box clocked it 2.4x faster. Output generation is 30%+ faster than Sonnet 5. It is also the first Sonnet to ship with the cyber safeguards Anthropic previously reserved for its top models.

Why the model tier stopped predicting quality for your business

The mental model most teams run — premium tier for hard work, cheap tier for bulk — just broke for a specific and common workload. For agentic coding, the $2 model beats the $4 model. That is not a rounding error you can ignore; on a pipeline doing real volume it is half your bill.

This is the second Anthropic launch in seven days. Opus 5.5 landed on September 22. If your model choice is hardcoded in three services and a cron job, you are one launch behind, permanently.

What to do this week:

Find every place a model string is written down. Not the ones you remember — grep for them. If claude-opus-5-5 or an older Sonnet appears anywhere you cannot change in under a minute, that is the actual problem, not the pricing.

Re-run your evals on Sonnet 5.5 before you switch anything. Terminal-Bench is Anthropic's workload. Your support triage, your invoice parsing, your product-copy generation are not. Twenty real cases, one afternoon.

Measure tool calls, not just tokens. The savings here come from the model doing less work, not costing less per unit. If your logging does not count tool calls per completed task, you will not see the 30% and you will not know if you lost it.

We build systems where swapping a model is a config change and a test run, not a sprint — because Anthropic ships on its schedule, not yours.

Key takeaways

  • Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0 vs Opus 5.5's 66.4%, at half the token price
  • Pricing is unchanged from Sonnet 5: $2 in / $10 out per million, $0.20 cache reads
  • The 30% cost saving comes from fewer tokens and fewer tool calls, not a lower sticker price
  • Available now as claude-sonnet-5-5 on Claude Platform, AWS, Google Cloud and Azure
  • Second Anthropic model launch in seven days — hardcoded model strings are now a recurring cost
  • Run your own evals; benchmark wins do not transfer to your prompts automatically

If finding every model string in your stack takes more than five minutes, that is the finding. We build AI automations where the model, the prompt and the spend ceiling are config, not code. Price what your AI work actually costs today, or send us your setup and we will tell you what to re-route.

Sources: Anthropic, Claude model docs, VentureBeat.

  • #claude-sonnet-5-5
  • #anthropic
  • #model-selection
  • #token-pricing
  • #ai-costs
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.