Skip to content
Rush Commerce
AI & Automation4 min read

Colossus 2 targets 1.21M Nvidia chips by December

Musk laid out a chip-by-chip schedule for Colossus 2: 550k now, 1.21 million by late December. What a compute glut does to your AI pricing.

Compute supply is usually a rumor. This week it came with a delivery schedule. Elon Musk posted the numbers on September 25: Colossus 2 runs 110,000 Nvidia GB200s and 440,000 GB300s today, with another 220,000 GB300s operational "next week," another 220,000 in November, and — "if we get lucky" — another 220,000 by late December. Add it up and the Memphis-area cluster lands near 1.21 million Nvidia chips by year-end, more than double where it sits now.

What actually happened

The post was a reply on X, not a press release, which is why it reads like a shipping manifest. Bloomberg picked it up the same day as the clearest timeline yet for the buildout at SpaceXAI, the entity formed when SpaceX absorbed xAI earlier this year.

For scale, Musk put Colossus 1 at 150k H100s, 50k H200s and 30k GB200s. Colossus 2's current 550,000 chips already dwarf it, and the GB300 is the newer, hungrier part.

The detail we keep coming back to is the increment. Everything moves in multiples of 110,000, and Musk's explanation is physical: it is the number of fiber optic cables that can be plugged into a central switch. Not a financing round, not a chip allocation — a cable count. Three of the four remaining tranches are scheduled; the last one is explicitly conditional.

Treat the year-end figure as a target from an interested party, not a delivered fact. Musk's own language — "if we get lucky" — is doing work. But the first 550,000 are running, and the November tranche is not speculative.

Why a compute buildout matters for your business

This is the supply side of your token price. You do not buy GB300s. You buy inference, and inference gets cheaper when capacity outruns demand. Frontier API prices have already fallen hard this year. A million-chip cluster coming online in a quarter is one more reason the number on your invoice next year should not resemble this year's.

So do not sign long. The single most expensive AI decision a small business made in 2026 was a two-year commit at 2026 prices. Keep contracts short, keep your prompts and eval harness portable, and re-shop the model layer quarterly. We build vendor-agnostic systems where the model is a config value, because the model you picked in March is rarely the right one in September.

Cheap compute is not free compute. Falling per-token prices have reliably produced rising bills, because teams answer a price cut with more calls, longer contexts and more agent loops. Instrument cost per completed job — a resolved ticket, an accepted draft — not cost per million tokens. The per-unit number can fall 60% while your monthly spend doubles.

The bottleneck is physical, and physical things slip. Cables, switches, transformers, substations, electricians. When your provider's capacity plan depends on how many fibers fit a switch, "next week" is a forecast. Design your automation so a capacity crunch degrades it instead of stopping it: queue non-urgent work, fall back to a smaller model, and never put a frontier API call on the critical path of a checkout.

Concentration is the real risk. Four or five companies are building nearly all of this. If your operations sit on one of them, you have inherited their permitting, power and politics. Keep a second provider wired and tested, even if it handles 5% of traffic.

Key takeaways

  • Musk said September 25 that Colossus 2 currently runs 110,000 Nvidia GB200s and 440,000 GB300s — about 550,000 chips
  • Another 220,000 GB300s go operational "next week," 220,000 more in November, and possibly 220,000 more by late December
  • That path reaches roughly 1.21 million Nvidia chips by year-end — treat the last tranche as conditional ("if we get lucky")
  • Expansion moves in 110,000-chip increments because that is how many fiber optic cables fit a central switch
  • Colossus 1, for comparison: 150k H100, 50k H200, 30k GB200
  • More capacity means continued downward pressure on inference prices — keep AI contracts short and the model layer swappable
  • Measure cost per completed job, not per million tokens; price cuts usually raise total spend
  • Build graceful degradation and keep a tested second provider — the constraints here are physical and schedules slip

When compute doubles every few months, the worst thing you can own is a long contract and a hard-coded model name. We build automation where the model is a config value and the fallback path is tested, so a price cut is a win instead of a migration. See how we build vendor-agnostic systems, or run the numbers on your current AI spend.

Sources: Elon Musk on X, Bloomberg.

  • #compute
  • #nvidia
  • #inference-costs
  • #vendor-risk
  • #ai-infrastructure
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.