Gemini 3.6 Flash: cheaper tokens, and 17% fewer of them
Google shipped Gemini 3.6 Flash at $1.50/$7.50 per million tokens and it burns 17% fewer output tokens. Two discounts stack — if your model layer is swappable.
Price cuts get the headline. Token efficiency is the one that actually moves your invoice. Google released Gemini 3.6 Flash today alongside 3.5 Flash-Lite and a security-tuned 3.5 Flash Cyber, and the interesting part isn't that the rate card went down. It's that the same job now costs fewer tokens and each token costs less. Those two discounts multiply.
What actually happened
Per Google's announcement, three models shipped July 21:
- Gemini 3.6 Flash — $1.50 per million input tokens, $7.50 per million output, down from $9 per million output on 3.5 Flash. Google says it uses 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index. On the DeepSWE coding benchmark it scores 49% versus 37% for its predecessor, per 9to5Google. Knowledge cutoff moved from January 2025 to March 2026. Available through the Gemini API, AI Studio, Android Studio, Antigravity, and Gemini Enterprise.
- Gemini 3.5 Flash-Lite — $0.30 / $2.50 per million, roughly 350 output tokens per second. The cheap, fast tier.
- Gemini 3.5 Flash Cyber — a vulnerability-finding model that is not generally available. It's a limited-access pilot for governments and trusted partners through CodeMender. Interesting, not purchasable.
DeepMind also said it has started pre-training Gemini 4. No date, no specifics.
Why Gemini 3.6 Flash matters for your business
If you run a workflow on 3.5 Flash — support triage, product description generation, document extraction — you just got a rate cut and a token cut on the same day. On a workload burning 40M output tokens a month, that's roughly $360 down to $250 on price alone, and the 17% token reduction takes it lower still.
You collect none of that if switching models means a code change, a redeploy, and a week of regression testing.
This is the argument we keep making and the news keeps making for us. The model is a component, not the architecture. Behind a provider boundary — one interface, model name in config, prompts and evals stored separately from the call site — a launch like today's is a Tuesday afternoon config change and a re-run of your eval set. Wired directly into your app logic, it's a project nobody schedules, and you keep paying last quarter's price for the next two years.
One honest caveat: benchmark numbers are not your workload. A 49% DeepSWE score tells you nothing about whether the model handles your invoice PDFs. The only number that matters is cost per completed task on your own data, measured before and after. If you don't have that baseline, you can't tell a price cut from a regression.
Key takeaways
- Gemini 3.6 Flash launched July 21 at $1.50/$7.50 per million tokens, down from $9 per million output on 3.5 Flash
- It also uses ~17% fewer output tokens per task — the rate cut and the efficiency gain compound
- 3.5 Flash-Lite lands at $0.30/$2.50; Flash Cyber is a limited pilot, not something you can buy
- You only capture the savings if the model name is config, not code — and only if you baseline cost per completed task
Paying last quarter's token price? We build AI features with the model behind a swappable boundary, so a price cut is a config change. Run the numbers on your workflow or tell us what you're running.
Sources: Google, 9to5Google.
- #gemini
- #llm
- #ai-costs
- #model-routing
Tommy Rush — Founder, Rush Commerce
Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More
Get The Rush Report weekly — one email, zero fluff.
Keep reading
Pilot Protocol's $4.5M: a directory isn't a standard
A seed-stage AI agent network wants to be where agents discover each other. Useful, but know the difference between an open spec and someone else's front door.
Read itNvidia backs Safe Superintelligence, discloses no terms
Nvidia's Safe Superintelligence partnership names no dollar figure and no term length. Reported at $5B. Here's how to read AI deals that omit the numbers.
Read it