Skip to content
Rush Commerce
AI & Automation3 min read

AMD's Helios ships with Microsoft: NVIDIA gets a rival

AMD launched its Helios rack and MI450 GPUs at Advancing AI 2026 with Microsoft Azure as anchor customer — the first credible rack-scale rival to NVIDIA.

For three years the answer to "who else makes AI infrastructure?" has been "nobody at scale." At Advancing AI 2026 in San Francisco this week, AMD tried to change that answer — and it brought Microsoft along as proof. The company launched its Helios rack-scale system and Instinct MI450 GPUs, with Azure signed on as anchor customer.

What actually happened

Per Yahoo Finance's coverage of the July 22–23 event, Helios is sold as one integrated unit — a full rack of Instinct MI455X accelerators paired with EPYC "Venice" CPUs (the first x86 server chip in volume production on TSMC's 2nm node), Pensando networking, and the ROCm software stack, delivering 31 TB of HBM4.

The parts that matter:

  • Microsoft is the anchor. Azure will deploy Helios to power frontier-model inference, with shipments beginning in the second half of 2026. Meta, OpenAI, Oracle, and TCS are named early adopters.
  • The gap is still enormous. AMD holds roughly 4.5% of the data-center GPU market against NVIDIA's ~95%, per Futurum Group estimates cited in the reporting. Racks are reportedly priced around $5 million each — not confirmed by AMD.
  • The runway is long. AMD and Anthropic also expanded their 2-gigawatt MI450 partnership, with AMD committing up to $5 billion in equity and first deployment starting in H1 2027, per the companies' announcement.

Why a second GPU supplier matters for your business

You don't buy racks. You buy tokens. But the price of your tokens is downstream of whether the people selling compute have any competition — and until this week, they didn't. A credible #2 with a hyperscaler anchor is the first real pressure on a market that's been a single-vendor toll road.

The catch: that pressure only reaches your invoice if your stack can actually move. NVIDIA's moat isn't the silicon, it's CUDA. AMD's answer is ROCm. If your inference is welded to one vendor's runtime, you'll watch the price war from the sidelines and pay list.

So the play is the boring one we keep repeating: keep your workloads chip-agnostic. Run inference through an abstraction that doesn't care whether it lands on CUDA or ROCm — the same logic behind free inference on any chip and keeping your software portable. The competition is finally arriving. Be built to take the cheaper price when it does.

Key takeaways

  • AMD launched Helios (72 MI455X GPUs, EPYC "Venice" on 2nm, 31 TB HBM4) and MI450 at Advancing AI 2026
  • Microsoft Azure is the anchor customer, with shipments starting H2 2026; Meta, OpenAI, Oracle, and TCS are early adopters
  • AMD still holds ~4.5% of the data-center GPU market vs NVIDIA's ~95% (Futurum)
  • A credible second supplier is the first real price pressure on compute in years
  • You only capture that pressure if your inference layer runs on CUDA or ROCm without a rewrite

Locked into one model vendor's runtime? We build inference layers that swap providers without a rewrite, so a price war works in your favor. See how we keep stacks portable or tell us what you're running.

Sources: Yahoo Finance, AMD/Anthropic via GlobeNewswire.

  • #amd
  • #nvidia
  • #inference
  • #compute
  • #portability
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.