Trainium gets Nvidia memory: your silicon hedge narrows
AWS will put Nvidia's NVLink Fusion and custom NVHBM into next-gen Trainium. The not-Nvidia option in your cost model is becoming less not-Nvidia.
Every AI cost model we've seen from a hopeful operator has the same escape hatch penciled in the margin: if Nvidia stays expensive, we move to the cheaper silicon. AWS just made that hatch smaller. Alongside a very large GPU order, AWS is putting Nvidia interconnect and Nvidia memory technology inside Trainium — the chip that was supposed to be the alternative.
What actually happened
Per Amazon's own announcement and Nvidia's newsroom release on August 26, AWS will deploy 2 million additional Nvidia GPUs across 2027–2028 — Blackwell Ultra, Rubin, and Rubin Ultra. That's on top of the 1 million-plus commitment announced at GTC 2026, which AWS says demand outran.
The part that changes the shape of the market is further down the page. AWS will support Nvidia NVLink Fusion in next-generation Trainium chips, and the two companies will integrate Nvidia's custom high-bandwidth memory (NVHBM) with Trainium for faster, more power-efficient memory. Nvidia Vera CPUs are also coming to AWS, aimed squarely at the CPU-side work behind agentic AI: code execution, tool use, sandboxing, data pipelines, orchestration. The companies also plan AI factories for the US government including 100,000 GPUs on AWS's secure infrastructure.
AWS cites its own figures for GPU-accelerated data processing — up to 3.7x faster with 30% better price-performance, and up to 9x faster vector indexing at a quarter of the cost. Vendor benchmarks; treat them as such until you run your workload.
Why it matters for your business
You are not buying chips. You are buying inference by the token, and the reason the silicon question ever mattered to you is that competition between vendors is what makes tokens get cheaper. When the leading alternative accelerator licenses the incumbent's interconnect and memory, that competition is now partly a supply relationship. Trainium may still be the cheaper line item. It's just no longer independent of Nvidia's roadmap or Nvidia's memory supply.
The practical read: stop treating chip diversity as your hedge. It was never a lever you controlled. Your actual portability lives one layer up, and it's the boring stuff — prompts, eval sets, and routing logic in your repo instead of a vendor console; an OpenAI-compatible interface in front of whatever serves your model; at least one workload you've genuinely run on open weights, so you know the switching cost by measurement rather than by assumption.
Then do the thing almost nobody does: put a date on the calendar to re-benchmark. Price-performance claims move every quarter now, and a hedge you never test is a paragraph in a doc, not an option.
Key takeaways
- AWS will deploy 2 million more Nvidia GPUs in 2027–2028 (Blackwell Ultra, Rubin, Rubin Ultra), on top of a 1M+ commitment that demand outran
- Next-generation Trainium will support Nvidia NVLink Fusion, and Nvidia's custom NVHBM memory will be integrated with Trainium
- Nvidia Vera CPUs are coming to AWS for agentic workloads — code execution, tool use, sandboxing, orchestration
- The 3.7x data-processing and 9x vector-indexing figures are AWS's own; benchmark your workload before budgeting on them
- Chip-level diversity isn't your hedge — portable prompts, evals, routing, and a tested open-weights fallback are
If your model vendor doubled its price next quarter, how long would switching take? We build AI systems where the prompts, evals, and routing live in your repository, so changing the model underneath is a config edit and a test run. See how we work.
Sources: About Amazon, NVIDIA Newsroom.
- #nvidia
- #aws
- #trainium
- #vendor-lock-in
- #inference-costs
Tommy Rush — Founder, Rush Commerce
Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More
Get The Rush Report weekly — one email, zero fluff.
Keep reading
Okta Agent SSO is GA: your AI agents get real logins
Okta shipped Agent SSO to all core SSO plans at no extra cost. AI agents get short-lived, governed tokens instead of pasted API keys. Here's what to do with it.
Read itNvidia's $12.9B Hugging Face bid: your registry picks a side
Nvidia is reportedly closing in on a $12.9B Hugging Face acquisition. If your deploy pulls weights from the Hub, your model registry now belongs to a chip vendor.
Read it