Skip to content
Rush Commerce
AI & Automation3 min read

Nvidia's $89B data center quarter: compute is revenue

Nvidia posted $96.2B in Q2 revenue and guided to $108B, while AWS ordered another 2 million GPUs. Why your token prices are not about to fall.

Nvidia reported fiscal Q2 2027 on August 26 and the Nvidia data center quarter number is the one to write down: $89.0 billion, up 117% year over year, inside $96.2 billion total revenue. Then it guided next quarter to $108 billion. If you have been budgeting on the assumption that inference gets cheaper every six months because supply catches up, this quarter says the opposite. Supply is sold.

What actually happened

Per Nvidia's results, revenue was $96.2 billion, up 106% year over year, with data center at $89.0 billion, up 117%. Non-GAAP gross margin held at 75.0%. Q3 guidance is $108.0 billion, plus or minus 2% — and Nvidia states it assumes no data center compute revenue from China in that outlook. Jensen Huang's framing on the call was blunt: "Now, compute is revenue."

The demand side came with a name attached. On the same call, Nvidia said AWS will deploy an additional 2 million GPUs — Blackwell Ultra, Rubin, and Rubin Ultra — plus Vera CPUs, some paired with Rubin and some standalone, rolling out from this quarter through fiscal 2029. TechCrunch notes that roughly triples the more-than-1-million-GPU commitment Amazon made five months ago. Neither company disclosed the price.

Why compute pricing matters for your business

Read the shape of it. A vendor at 75% gross margin, guiding up 12% sequentially, with a top customer tripling a multi-year order, is not a vendor under pressure to cut prices. The capacity AWS just bought lands in 2027 and 2028. That is the window in which your inference bill gets negotiated, and it is already spoken for.

So stop modelling a price cut you do not control. Model the two things you do: how many tokens a finished task takes, and whether the task needs a frontier model at all. We have watched small operators cut AI spend 40-60% without touching a vendor contract — by routing the cheap 80% of calls to a smaller model, caching the prompts that repeat, and killing the retry loops nobody was measuring. That is engineering, and it survives whatever the rate card does.

The second move is portability. Keep the model name in config. Keep prompts and evals in version control, separate from your call sites. When a price changes or a cheaper open-weights model clears your bar, swapping should be a config change and an eval run, not a rebuild.

Key takeaways

  • Nvidia Q2 FY2027: $96.2B revenue (+106%), data center $89.0B (+117%), 75.0% non-GAAP gross margin
  • Q3 guidance is $108.0B ±2%, explicitly assuming zero data center compute revenue from China
  • AWS is deploying an additional 2 million GPUs plus Vera CPUs through fiscal 2029 — roughly triple its March commitment
  • Nobody in this supply chain has a reason to lower your token price before 2028
  • Control what you can: tokens per completed task, model routing, prompt caching, and a swappable model layer

Your AI cost line is an engineering problem, not a procurement one. We build routing, caching, and eval layers that cut token spend without locking you to one vendor. Run the numbers or tell us what you're spending.

Sources: Nvidia, TechCrunch.

  • #nvidia
  • #ai-costs
  • #compute
  • #gpu
  • #vendor-pricing
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.