Gemini 3.7 Flash is half price until January 1
Gemini 3.7 Flash launched at $0.75/$3.75 per million tokens. On January 1, 2027 it doubles. A discount with a published expiry is a budget event, not a price cut.
Google shipped Gemini 3.7 Flash on August 13 — three weeks after 3.6 Flash — at half the price. It also published the date the discount ends. That second part is the news. An introductory rate with January 1, 2027 stamped on it isn't a price cut; it's a price increase with a grace period, and you now have about 140 days to find out whether your workload survives it.
What actually happened
Per Google's announcement, Gemini 3.7 Flash runs at $0.75 per million input tokens and $3.75 per million output through December 31, 2026. On January 1, 2027 it goes to $1.50 / $7.50 — which is exactly what 3.6 Flash cost at launch. The intro rate is a discount off the real number, not a new floor.
The capability gains are real and mostly aimed at coding and agents. Against 3.6 Flash: DeepSWE v1.1 65.3% vs 49.0%, FrontierCode 1.1 43.6% vs 34.4%, WebDev Arena 1588 Elo vs 1538, GDP.pdf document reasoning 34.0% vs 22.0%, and AutomationBench for enterprise workflow completion 30.4% vs 17.0%. Google credits developer feedback and algorithmic changes to the reasoning core rather than a fresh pretraining run, as VentureBeat reported. It's live in the Gemini API via AI Studio, Android Studio, Antigravity, and Gemini Enterprise.
Why the price cliff matters for your business
Run the arithmetic before you build on the intro rate. A workflow burning 30M output tokens a month costs about $113 today and about $225 in January. If your unit economics only clear at $0.75, you have a feature with an expiration date and nobody has told your P&L.
So put it on the calendar as a budget line, not a footnote. Then do the thing that makes the date survivable: keep the model name in config, keep prompts and evals in version control separate from your call sites, and measure cost per completed task — not cost per token — on your own data. When the rate resets, the question "does a cheaper model still pass our evals?" should take an afternoon to answer, not a quarter. That's the same swappable model layer we've argued for all year, and the release cadence keeps proving the point: three weeks between Flash versions is shorter than most teams' procurement cycle.
One honest caveat. A DeepSWE score of 65.3% says nothing about whether this model reads your purchase orders correctly. Benchmarks are a reason to run an eval, not a reason to switch.
Key takeaways
- Gemini 3.7 Flash launched August 13 at $0.75/$3.75 per million tokens — half of 3.6 Flash's launch price
- The introductory rate expires December 31, 2026; on January 1, 2027 it doubles to $1.50/$7.50
- Big coding and agent gains over 3.6 Flash: DeepSWE 65.3% vs 49.0%, AutomationBench 30.4% vs 17.0%
- Three weeks between Flash releases — treat model choice as config you can re-evaluate, not architecture you commit to
- Size your workload at the post-January price and baseline cost per completed task, not cost per token
Does your AI feature still pencil at double the token price? Model what the January reset does to your margin, then let us build the layer that makes swapping models a config change. Run the numbers or tell us what you're running.
Sources: Google, VentureBeat.
- #gemini
- #ai-costs
- #llm
- #model-routing
Tommy Rush — Founder, Rush Commerce
Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More
Get The Rush Report weekly — one email, zero fluff.
Keep reading
Same model, 41% cheaper: the harness is the cost lever
Writer's research cut agent cost per task from $0.21 to $0.12 across six foundation models by changing only the orchestration layer. Token costs are an engineering problem.
Read itVantage eyes a $100B IPO: your AI landlord goes public
Reuters reports Vantage Data Centers is exploring a $100B IPO or sale. When the compute landlord answers to public markets, your token price gets a quarterly cadence.
Read it