Skip to content
Rush Commerce
AI & Automation3 min read

Gemini 3.7 Flash is half price until January 1

Gemini 3.7 Flash launched at $0.75/$3.75 per million tokens. On January 1, 2027 it doubles. A discount with a published expiry is a budget event, not a price cut.

Google shipped Gemini 3.7 Flash on August 13 — three weeks after 3.6 Flash — at half the price. It also published the date the discount ends. That second part is the news. An introductory rate with January 1, 2027 stamped on it isn't a price cut; it's a price increase with a grace period, and you now have about 140 days to find out whether your workload survives it.

What actually happened

Per Google's announcement, Gemini 3.7 Flash runs at $0.75 per million input tokens and $3.75 per million output through December 31, 2026. On January 1, 2027 it goes to $1.50 / $7.50 — which is exactly what 3.6 Flash cost at launch. The intro rate is a discount off the real number, not a new floor.

The capability gains are real and mostly aimed at coding and agents. Against 3.6 Flash: DeepSWE v1.1 65.3% vs 49.0%, FrontierCode 1.1 43.6% vs 34.4%, WebDev Arena 1588 Elo vs 1538, GDP.pdf document reasoning 34.0% vs 22.0%, and AutomationBench for enterprise workflow completion 30.4% vs 17.0%. Google credits developer feedback and algorithmic changes to the reasoning core rather than a fresh pretraining run, as VentureBeat reported. It's live in the Gemini API via AI Studio, Android Studio, Antigravity, and Gemini Enterprise.

Why the price cliff matters for your business

Run the arithmetic before you build on the intro rate. A workflow burning 30M output tokens a month costs about $113 today and about $225 in January. If your unit economics only clear at $0.75, you have a feature with an expiration date and nobody has told your P&L.

So put it on the calendar as a budget line, not a footnote. Then do the thing that makes the date survivable: keep the model name in config, keep prompts and evals in version control separate from your call sites, and measure cost per completed task — not cost per token — on your own data. When the rate resets, the question "does a cheaper model still pass our evals?" should take an afternoon to answer, not a quarter. That's the same swappable model layer we've argued for all year, and the release cadence keeps proving the point: three weeks between Flash versions is shorter than most teams' procurement cycle.

One honest caveat. A DeepSWE score of 65.3% says nothing about whether this model reads your purchase orders correctly. Benchmarks are a reason to run an eval, not a reason to switch.

Key takeaways

  • Gemini 3.7 Flash launched August 13 at $0.75/$3.75 per million tokens — half of 3.6 Flash's launch price
  • The introductory rate expires December 31, 2026; on January 1, 2027 it doubles to $1.50/$7.50
  • Big coding and agent gains over 3.6 Flash: DeepSWE 65.3% vs 49.0%, AutomationBench 30.4% vs 17.0%
  • Three weeks between Flash releases — treat model choice as config you can re-evaluate, not architecture you commit to
  • Size your workload at the post-January price and baseline cost per completed task, not cost per token

Does your AI feature still pencil at double the token price? Model what the January reset does to your margin, then let us build the layer that makes swapping models a config change. Run the numbers or tell us what you're running.

Sources: Google, VentureBeat.

  • #gemini
  • #ai-costs
  • #llm
  • #model-routing
  • #google
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.