Positron raises $875M for inference chips without HBM
Positron AI raised $875M at a $5B valuation for inference silicon built on commodity LPDDR5X instead of HBM. Why the memory choice decides your token price.
Positron AI raised $875 million at a $5 billion post-money valuation to ship inference silicon that deliberately skips HBM. That last part is the story. Everything you pay per token traces back to a memory decision made in a fab two years before you swipe the card.
What actually happened
Positron announced the round on September 10 in two tranches: a $375 million Series C priced at a $3.5 billion pre-money valuation, co-led by NEA, Andra Capital, Atreides Management, Valor Equity Partners, and Dylan Patel's SemiAnalysis Capital, plus a $500 million Series C-1 led by NEA and Netscape co-founder Jim Clark.
The technical claim is specific. Positron builds around memory capacity and bandwidth rather than raw FLOPS, and says its systems hit over 90% memory bandwidth utilization using commodity LPDDR5X — which sidesteps the HBM and CoWoS packaging bottlenecks that gate everyone else's supply. Atlas, the shipping first-generation system, is deployed at Oracle Cloud Infrastructure across 50-plus racks, with production customers including Parasail, Jump Trading, and i3d.net.
Next up: Asimov, custom silicon on TSMC N3P, taping out late 2026 with production expected in the second half of 2027, offering 288 GB to 2,304 GB of memory per chip. Titan packs four to eight Asimov chips into a node aimed at models past 16 trillion parameters and context windows beyond 10 million tokens.
Read the timeline honestly. Asimov is a 2027 product. The money is real today; the silicon that justifies it is not.
Why inference memory economics matter for your business
You do not buy chips. You buy tokens, and the price of a token is mostly a memory bill.
Generation is memory-bandwidth bound, not compute bound. Every token you stream out re-reads the model weights and the KV cache. That is why long context costs more than long output, why cache hits are the single biggest lever on an agent's bill, and why HBM scarcity shows up in your invoice months later as a price increase or a rate limit. Chinese chipmakers are reportedly raising prices on HBM shortages right now. Same physics, different continent.
So the practical read is not "switch to Positron" — you cannot, they sell systems to clouds, not to you. The read is that a second supply chain for inference memory is getting funded, and a second supply chain is what breaks a price floor. If you have been assuming your per-million-token cost only goes down, that assumption is finally getting some structural support behind it, arriving in 2027.
What you do this quarter is unglamorous. Instrument your cache hit rate. Measure prompt tokens versus output tokens per workflow, because they price differently and scale differently. Keep your agent code behind a model-agnostic interface so a cheaper provider is a config change and not a rewrite. The people who benefit from a token price war are the ones who can switch the day it starts.
Key takeaways
- Positron AI raised $875M at a $5B post-money valuation: a $375M Series C plus a $500M Series C-1 led by NEA and Jim Clark
- Its systems use commodity LPDDR5X instead of HBM, avoiding HBM and CoWoS supply constraints, at 90%+ memory bandwidth utilization
- Atlas ships today at Oracle Cloud Infrastructure across 50+ racks; customers include Parasail, Jump Trading, i3d.net
- Asimov silicon tapes out on TSMC N3P in late 2026 with production in 2H 2027 — the payoff is a year and a half out
- Token generation is memory-bandwidth bound, so memory supply, not compute, sets your inference price floor
- Instrument cache hit rate and prompt-vs-output token mix now, and keep model calls behind a swappable interface
Cheap tokens only help you if you can reach them. We build AI systems with the provider behind an interface, cache behavior you can measure, and cost telemetry per workflow — so a price cut is something you capture instead of read about. Run the numbers on your automation spend, or tell us what your agent bill actually looks like.
Sources: Positron AI press release (PR Newswire), Quartz.
- #inference
- #ai-chips
- #token-costs
- #positron
- #vendor-strategy
Tommy Rush — Founder, Rush Commerce
Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More
Get The Rush Report weekly — one email, zero fluff.
Keep reading
Zscaler Agentic SOC: the agent now pulls the trigger
Zscaler's Agentic SOC uses Anthropic and OpenAI models to triage, investigate, and contain threats automatically. The line that moved is autonomous remediation.
Read itListen Labs scrapped $125M for Salesforce talks
Listen Labs walked away from a signed $125M term sheet to talk acquisition with Salesforce. What it means when your AI vendor becomes a SKU.
Read it