Claude Haiku 5.5: 90% cheaper under 100K tokens
Claude Haiku 5.5 costs $0.10/$0.50 per million tokens under 100K-token prompts, 90% below Haiku 4.5. Here is how to rerun your AI bill before you switch.
The cheap model just got a lot cheaper. Anthropic released Claude Haiku 5.5 today at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. That is 90% below Haiku 4.5, and it matches OpenAI's GPT-6 Luna to the cent. If your business runs ticket triage, product tagging, or email summaries on an API, your bill for that work could drop by most of a zero. Read the fine print first.
What actually happened
Per Anthropic's announcement, Haiku 5.5 has two price tiers. Prompts up to 100K tokens cost $0.10 input / $0.50 output per million tokens. Prompts over 100K cost $0.50 / $2.50. Haiku 4.5 was $1.00 / $5.00 flat. Cache reads drop to $0.01 per million in the lower tier.
Anthropic says the average cost to run is about 75% lower, not 90%. Two reasons: roughly 90% of Haiku 4.5 requests fell under 100K tokens, and the new tokenizer "uses slightly more tokens per task." VentureBeat lays the price against the field: GPT-6 Luna at $0.10 / $0.50, Gemini 3.5 Flash-Lite at $0.30 / $2.50.
Other details that matter:
- It is the first Haiku with an adjustable effort setting. The default is medium.
- Model ID
claude-haiku-5-5, live on Anthropic's API, AWS, Google Cloud, and Azure. - Anthropic recommends it for summaries, classification, database queries, live support, and as a subagent under Sonnet 5.5 or Opus 5.5. For complex agentic coding, it points you to the bigger models.
- Sonnet 5.5 cache reads were cut from $0.20 to $0.10 per million tokens.
- Max and Team subscribers get monthly API credits: $100 (Max 5x), $200 (Max 20x), up to $500 pooled for Team.
Why it matters for your business
Do the math on your own traffic. Take a classification job with 2,000 input tokens and 200 output tokens per request, 100,000 requests a month. On Haiku 4.5 that is about $300. On Haiku 5.5 it is about $30, plus whatever the new tokenizer adds. Small numbers, but the same ratio holds at ten times the volume.
Watch the 100K line. If you stuff whole catalogs, contracts, or long chat histories into each prompt, you pay the upper tier. That is still 50% cheaper than Haiku 4.5, not 90%. Trimming context is now worth real money.
Test before you swap. Vendor benchmarks are vendor benchmarks. Run your last 500 real requests through both models, compare outputs, and check the token counts. A price cut that changes your outputs is not a saving.
Jobs you priced out are back in. Tagging every product review, summarizing every support call, scoring every inbound lead. At $0.10 per million input tokens, the model cost on most small-business volumes stops being the line item that kills the idea.
Key takeaways
- Claude Haiku 5.5 costs $0.10 / $0.50 per million tokens for prompts up to 100K tokens
- Anthropic's own average estimate is about 75% cheaper than Haiku 4.5, because of the tier split and a new tokenizer
- Prompts over 100K tokens cost $0.50 / $2.50, so long-context jobs save 50%, not 90%
- Haiku 5.5 is now priced the same as GPT-6 Luna
- Rerun a sample of real traffic before you switch; check outputs and token counts
Not sure what your AI workload actually costs? We build automations with the model choice in config, not hardcoded, so a price cut like this one is a one-line change. Put your volumes into our ROI calculator, or send us your current setup and we will price the switch.
Sources: Anthropic, VentureBeat.
- #claude
- #haiku-5-5
- #anthropic
- #llm-pricing
- #ai-costs
Tommy Rush — Founder, Rush Commerce
Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More
Get The Rush Report weekly — one email, zero fluff.
Keep reading
nCino + Parlay: AI SBA loan intake comes to 1,400 lenders
nCino will resell Parlay's AI SBA loan intake to banks. If you plan to borrow, an AI now screens your file first. Here's how to make it credit-ready.
Read itGemini coworker agents get email and Drive: onboard them like hires
Google's Gemini at Work 2026 gives coworker agents their own email, calendar and Drive. Treat each one like a new hire: scoped access, a budget, an audit log.
Read it