Skip to content
Rush Commerce
AI & Automation3 min read

Google's Frozen v2 chip bakes Gemini into the silicon

Google is reportedly building Frozen v2, a chip with Gemini's architecture etched in, at 6-10x tokens per watt. Why cheap inference is getting model-specific.

The name is the whole story. Google is reportedly developing a server chip called Frozen v2 — frozen because parts of Gemini's architecture would be permanently etched into the circuitry. The efficiency claim is real money: six to ten times the tokens per watt versus Google's newest TPUs. The catch is in the same sentence. A chip that only runs one model architecture only stays cheap while that architecture stays frozen.

What actually happened

The Information broke the story on July 20. Per CNBC, Google engineers project Frozen v2 could serve between six and ten times more tokens per unit of power than the company's latest TPUs, with rollout to Google data centers targeted as soon as 2028. Alphabet shares closed up about 1.5%.

The engineering is straightforward once you see it. Hardwiring the model's structure removes runtime overhead — fewer calculations per response, less data shuttled between memory and compute. SiliconANGLE reports the design includes enough onboard memory to run Gemini without spilling to external RAM, which is the bottleneck that eats real-world inference throughput.

The constraint is equally straightforward: the chip works with future Gemini models only if Google keeps the same underlying architecture. Google reportedly treats Frozen v2 as a trial run rather than a TPU-scale program, and hasn't confirmed the project. This is a report about 2028, not a product page. Treat it as a signal, not a plan.

Why model-specific silicon matters for your business

We've written about transformer-only ASICs setting the floor under your token bill. Frozen v2 is a step further down that road, and the direction matters more than the chip.

Generic inference silicon makes all models cheaper, and cheap spreads across vendors. Model-specific silicon makes one model cheaper — and that discount lives inside one vendor's stack. If the cheapest tokens in 2028 are Gemini tokens on Gemini hardware, the price gap becomes a gravity well. That's not a conspiracy; it's just what happens when the economics stop being portable.

There's a second-order effect worth noting: it constrains Google too. You cannot casually redesign an architecture you've committed to silicon on a multi-year fab timeline. Frozen cuts both ways.

Nothing here changes what you should do this quarter, which is the point. Keep the model a configuration value, not a dependency baked through your codebase. Measure cost per completed task, not cost per token, so you can tell a real price drop from a repricing. Keep at least one alternative provider wired and tested, even if you never route to it. When a vendor's price advantage comes from hardware you can't buy, portability is the only leverage you have left.

Key takeaways

  • The Information reported July 20 that Google is developing "Frozen v2," a chip with Gemini's architecture etched into the silicon
  • Engineers project 6–10x more tokens per watt than Google's latest TPUs, with data-center rollout targeted for 2028
  • It reportedly works only with future Gemini models that keep the same architecture; Google hasn't confirmed the project
  • Model-specific silicon concentrates the price advantage inside one vendor's stack instead of spreading it across the market
  • Keep the model a config value, track cost per completed task, and keep a second provider wired and tested

One model provider deep in your codebase? We build AI systems where the model is a setting and switching providers is a one-line change — so a price drop is something you capture, not something you read about. See what we build.

Sources: CNBC — Alphabet stock pops on report it's developing a more efficient AI chip, SiliconANGLE.

  • #google
  • #gemini
  • #ai-chips
  • #inference
  • #vendor-lock-in
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.