Cerebras CS-5 hits 2027. Buy the speed that ships now
Cerebras laid out a wafer-scale roadmap at Hot Chips 2026: CS-5 in 2027, CS-6 with stacked DRAM later. Don't design your product around unpurchasable latency.
Cerebras used Hot Chips 2026 to publish a wafer-scale roadmap that runs two generations past what you can buy. The headline number — 10,000 output tokens per second per user — belongs to a system arriving in 2027. The number you can actually put in front of a customer this quarter is the one already racked.
What actually happened
Per Cerebras' own Hot Chips deep dive, the shipping system is CS-4 on the Nexus rack platform: three wafer-scale engines in modular compute backpacks, with AC/DC converters sitting 0.5mm from the wafer instead of the roughly 50mm typical in GPU systems. That proximity is why they can push nearly twice the power at almost the same voltage. ServeTheHome's session notes put CS-4 at about 2x the tokens of CS-3 at 10x the tokens per watt, with backpacks delivering double the power and cooling using 50% fewer components.
Then the roadmap. CS-5, targeted for 2027: up to 10,000 output tokens per second per user on open models, up to 5,000 on frontier models, 3 million tokens per second per megawatt, and support for models past 50 trillion parameters. CS-6, further out with no date: DRAM stacked directly on the wafer for the first time, trading on-chip SRAM for compute density and letting the stacked memory carry capacity. Tom's Hardware covered the stacking claim as the structural bet of the roadmap.
Nexus itself is pitched as a platform designed to double token-generation speed year over year. That is a design goal in a slide deck, not a purchase order.
Why the inference speed roadmap matters for your business
Latency is a product decision, and roadmaps don't ship features. If a workflow only works at 10,000 tokens per second, it does not work in 2026. Design against the throughput you can buy today — we covered what today's dial actually looks like — and treat next year's number as upside you get for free if it arrives.
Read every headline figure back to its footnote. Cerebras separates open-model and frontier-model rates for a reason: the frontier number is half. Per-user token rates, per-megawatt efficiency, and aggregate throughput are three different claims, and vendors mix them in the same paragraph. We flagged the same pattern in the CS-4 sparsity numbers.
Speed this specific is a lock-in risk. A UX that depends on one vendor's wafer-scale latency is a UX you cannot move. Keep the fast path behind an interface you own, so a cheaper provider next year is a config change instead of a redesign.
Key takeaways
- Cerebras CS-4 on Nexus is the shipping system: ~2x the tokens of CS-3 at 10x tokens per watt
- CS-5 targets 2027 — up to 10,000 output tokens/sec per user on open models, 5,000 on frontier models
- CS-5 also claims 3M tokens/sec per megawatt and support for 50T+ parameter models
- CS-6 plans 3D-stacked DRAM on the wafer, cutting SRAM for compute density. No announced date
- The frontier-model rate is half the open-model rate — check which number a vendor quoted you
- Build the fast path behind your own interface so vendor latency stays swappable
If your feature only works at a token rate that ships next year, you don't have a feature. We size AI workflows against latency you can buy this quarter, then keep the inference provider behind a routing layer you own. See how we build vendor-agnostic inference, or bring us the workflow and we'll tell you if the speed is there yet.
Sources: Cerebras, ServeTheHome, Tom's Hardware.
- #inference-speed
- #cerebras
- #ai-infrastructure
- #latency
- #vendor-roadmaps
Tommy Rush — Founder, Rush Commerce
Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More
Get The Rush Report weekly — one email, zero fluff.
Keep reading
Open-weight models: 6.1% adoption, billions in deals
Ramp puts open-source model platforms at 6.1% of AI-using businesses while Nvidia and Stripe buy the layer. Why you probably should not self-host yet.
Read itNvidia pauses its AI cloud revenue-share program
Nvidia pulled a July financing program for AI clouds inside two months over antitrust and control concerns. Your inference vendor's credit line just changed.
Read it