Claude Sonnet 5.5 beats Opus 5.5 on coding at half the price
Claude Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0 against Opus 5.5's 66.4%, at $2/$10 per million tokens. Your default model just changed.
Anthropic shipped Claude Sonnet 5.5 today, and the number that matters is not the price — it is that the cheap model now outscores the expensive one on agentic coding. Sonnet 5.5 hits 70.6% on Terminal-Bench 4.0. Opus 5.5, which came out six days ago at twice the token price, hits 66.4%. If you routed your coding agents to Opus last week on the assumption that bigger is better, that assumption is now wrong and expensive.
What actually happened with Claude Sonnet 5.5
Anthropic released Sonnet 5.5 as the second model in the Claude 5.5 family, at $2 per million input tokens and $10 per million output — identical to Sonnet 5's pricing. Cache reads are $0.20 per million. It is available on the Claude Platform, AWS, Google Cloud and Microsoft Azure as claude-sonnet-5-5, with a 1M-token context window and 128K max output per the model docs.
The benchmark jumps from Sonnet 5 are not incremental. Terminal-Bench 4.0 goes from 10.3% to 70.6%. OSWorld 2.1, the computer-use test, from 57.0% to 80.1%. CursorBench 4.0 from 34.1% to 55.5%. GDPval-AA from 1449 to 1844 — Opus 5.5 scores 1846 on the same test. On most evals Sonnet 5.5 now lands within a point or two of Opus 5.5.
Anthropic's cost claim is subtler than a price cut: same sticker price, up to 30% less per task, because the model finishes with fewer tokens and fewer tool calls. Slack measured about 14% fewer output tokens. VentureBeat reports Lovable saw a third fewer tool calls and roughly half as many shell executions, and Box clocked it 2.4x faster. Output generation is 30%+ faster than Sonnet 5. It is also the first Sonnet to ship with the cyber safeguards Anthropic previously reserved for its top models.
Why the model tier stopped predicting quality for your business
The mental model most teams run — premium tier for hard work, cheap tier for bulk — just broke for a specific and common workload. For agentic coding, the $2 model beats the $4 model. That is not a rounding error you can ignore; on a pipeline doing real volume it is half your bill.
This is the second Anthropic launch in seven days. Opus 5.5 landed on September 22. If your model choice is hardcoded in three services and a cron job, you are one launch behind, permanently.
What to do this week:
Find every place a model string is written down. Not the ones you remember — grep for them. If claude-opus-5-5 or an older Sonnet appears anywhere you cannot change in under a minute, that is the actual problem, not the pricing.
Re-run your evals on Sonnet 5.5 before you switch anything. Terminal-Bench is Anthropic's workload. Your support triage, your invoice parsing, your product-copy generation are not. Twenty real cases, one afternoon.
Measure tool calls, not just tokens. The savings here come from the model doing less work, not costing less per unit. If your logging does not count tool calls per completed task, you will not see the 30% and you will not know if you lost it.
We build systems where swapping a model is a config change and a test run, not a sprint — because Anthropic ships on its schedule, not yours.
Key takeaways
- Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0 vs Opus 5.5's 66.4%, at half the token price
- Pricing is unchanged from Sonnet 5: $2 in / $10 out per million, $0.20 cache reads
- The 30% cost saving comes from fewer tokens and fewer tool calls, not a lower sticker price
- Available now as
claude-sonnet-5-5on Claude Platform, AWS, Google Cloud and Azure - Second Anthropic model launch in seven days — hardcoded model strings are now a recurring cost
- Run your own evals; benchmark wins do not transfer to your prompts automatically
If finding every model string in your stack takes more than five minutes, that is the finding. We build AI automations where the model, the prompt and the spend ceiling are config, not code. Price what your AI work actually costs today, or send us your setup and we will tell you what to re-route.
Sources: Anthropic, Claude model docs, VentureBeat.
- #claude-sonnet-5-5
- #anthropic
- #model-selection
- #token-pricing
- #ai-costs
Tommy Rush — Founder, Rush Commerce
Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More
Get The Rush Report weekly — one email, zero fluff.
Keep reading
x47.c botnet drains your AI API credits with stolen keys
A $200 Windows botnet ships an AI API drain module that burns your OpenAI credits with your own key. Your site stays up while the bill lands.
Read itNvidia puts AI agent containment in the silicon
Nvidia's Open Agent Safety Platform pairs open-source OpenShell with a Sentry watchdog on BlueField-4 DPUs that kills out-of-bounds agents in milliseconds.
Read it