Skip to content
Rush Commerce
AI & Automation3 min read

Cerebras CS-5 hits 2027. Buy the speed that ships now

Cerebras laid out a wafer-scale roadmap at Hot Chips 2026: CS-5 in 2027, CS-6 with stacked DRAM later. Don't design your product around unpurchasable latency.

Cerebras used Hot Chips 2026 to publish a wafer-scale roadmap that runs two generations past what you can buy. The headline number — 10,000 output tokens per second per user — belongs to a system arriving in 2027. The number you can actually put in front of a customer this quarter is the one already racked.

What actually happened

Per Cerebras' own Hot Chips deep dive, the shipping system is CS-4 on the Nexus rack platform: three wafer-scale engines in modular compute backpacks, with AC/DC converters sitting 0.5mm from the wafer instead of the roughly 50mm typical in GPU systems. That proximity is why they can push nearly twice the power at almost the same voltage. ServeTheHome's session notes put CS-4 at about 2x the tokens of CS-3 at 10x the tokens per watt, with backpacks delivering double the power and cooling using 50% fewer components.

Then the roadmap. CS-5, targeted for 2027: up to 10,000 output tokens per second per user on open models, up to 5,000 on frontier models, 3 million tokens per second per megawatt, and support for models past 50 trillion parameters. CS-6, further out with no date: DRAM stacked directly on the wafer for the first time, trading on-chip SRAM for compute density and letting the stacked memory carry capacity. Tom's Hardware covered the stacking claim as the structural bet of the roadmap.

Nexus itself is pitched as a platform designed to double token-generation speed year over year. That is a design goal in a slide deck, not a purchase order.

Why the inference speed roadmap matters for your business

Latency is a product decision, and roadmaps don't ship features. If a workflow only works at 10,000 tokens per second, it does not work in 2026. Design against the throughput you can buy today — we covered what today's dial actually looks like — and treat next year's number as upside you get for free if it arrives.

Read every headline figure back to its footnote. Cerebras separates open-model and frontier-model rates for a reason: the frontier number is half. Per-user token rates, per-megawatt efficiency, and aggregate throughput are three different claims, and vendors mix them in the same paragraph. We flagged the same pattern in the CS-4 sparsity numbers.

Speed this specific is a lock-in risk. A UX that depends on one vendor's wafer-scale latency is a UX you cannot move. Keep the fast path behind an interface you own, so a cheaper provider next year is a config change instead of a redesign.

Key takeaways

  • Cerebras CS-4 on Nexus is the shipping system: ~2x the tokens of CS-3 at 10x tokens per watt
  • CS-5 targets 2027 — up to 10,000 output tokens/sec per user on open models, 5,000 on frontier models
  • CS-5 also claims 3M tokens/sec per megawatt and support for 50T+ parameter models
  • CS-6 plans 3D-stacked DRAM on the wafer, cutting SRAM for compute density. No announced date
  • The frontier-model rate is half the open-model rate — check which number a vendor quoted you
  • Build the fast path behind your own interface so vendor latency stays swappable

If your feature only works at a token rate that ships next year, you don't have a feature. We size AI workflows against latency you can buy this quarter, then keep the inference provider behind a routing layer you own. See how we build vendor-agnostic inference, or bring us the workflow and we'll tell you if the speed is there yet.

Sources: Cerebras, ServeTheHome, Tom's Hardware.

  • #inference-speed
  • #cerebras
  • #ai-infrastructure
  • #latency
  • #vendor-roadmaps
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.