Chinese AI chips jump up to 50% on HBM shortage
Huawei and Cambricon raised AI accelerator prices 20-50% as high-bandwidth memory tightens. Memory, not logic, is the floor under your inference bill.
Huawei and Cambricon have raised AI accelerator prices by 20% to 50% in two months, and the cause is memory. Reuters reported on September 10 that a worldwide high-bandwidth memory shortage is pushing up the cost of Chinese AI silicon faster than anything happening on the logic side. It is the clearest public price signal yet that the constraint in AI hardware moved off the GPU die and onto the memory stacked next to it.
What actually happened
Huawei repriced the Ascend 950DT accelerator card above 250,000 yuan — about $37,255 — a jump of 20% to 50% over quotes given to customers two months earlier, depending on contract terms. The Ascend 950PR went from roughly 60,000 yuan to more than 80,000. The older 910C climbed from about 90,000 yuan to more than 110,000.
Cambricon repriced its next-generation part, tentatively the 690, at 20% to 30% above its earlier indications. MetaX and Iluvatar CoreX are in the same market with the same input costs.
The mechanism is straightforward. HBM is supplied almost entirely by SK Hynix, Samsung, and Micron. Washington tightened export controls on advanced HBM to China in December 2024, so Chinese buyers work through grey-market channels at several times the price the rest of the world pays. Memory is a large share of an accelerator's bill of materials, so that cost lands directly on the finished card.
Why an AI hardware shortage matters for your business
Read past the China framing. Export controls explain why the increase shows up first and worst in Chinese pricing. They do not explain the shortage, which is global. Everyone building inference capacity is bidding for the same HBM out of the same three suppliers, and that is upstream of every token you buy.
We covered Positron's $875 million round for inference silicon that routes around HBM two days ago. That is a $5 billion bet on exactly this constraint. When a component shortage is severe enough to fund an architectural detour, it is not a quarter-long blip.
So do the unglamorous work now. Instrument tokens per completed task — not per request, per task — so you know what a 20% input-cost increase does to your margin before your vendor tells you. Turn on prompt caching if your provider offers it; most teams leave a large discount on the table because nobody checked. Look hard at whether your agent needs a frontier model for every hop, because a lot of orchestration burns premium tokens on "yes" and "confirm the order."
And keep an open-weights path warm. Not because you want to run inference yourself — you almost certainly don't — but because a 7B model on commodity hardware handling your classification and routing steps is a real hedge, and it takes a weekend to prove out. The teams that get squeezed by hardware pricing are the ones with a single provider and no measurement.
Key takeaways
- Huawei's Ascend 950DT is now quoted above 250,000 yuan (~$37,255), up 20-50% from two months ago
- Cambricon's next-gen 690 was repriced 20-30% higher; the Ascend 950PR and 910C also rose
- The driver is a global HBM shortage, made worse in China by December 2024 export controls and grey-market pricing
- Memory is a large share of accelerator cost, so the increase passes straight through to inference capacity
- Measure tokens per completed task and enable prompt caching before your provider's costs reach your invoice
- Keep a small open-weights model validated for classification and routing steps as a real hedge
Most AI bills are 60% waste and nobody's measured it. We instrument cost per completed task, cut the premium-model calls that don't need to be premium, and leave you with numbers you can defend. Model your AI spend, or send us your current invoice and we'll tell you what's burnable.
- #ai-hardware
- #hbm
- #inference-costs
- #supply-chain
- #vendor-pricing
Tommy Rush — Founder, Rush Commerce
Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More
Get The Rush Report weekly — one email, zero fluff.
Keep reading
AI spend per employee fell 10% at the top spenders
Ramp's August AI Index shows median AI spend per employee down nearly 10% while token prices fell 41% since March. Re-run your automation cost math.
Read itGoogle's €13B Finland bet runs on a 22-year nuclear PPA
Google is spending €13 billion on four Finnish data center sites and signed a 22-year Loviisa nuclear deal. What a two-decade power contract says about AI pricing.
Read it