Skip to content
Rush Commerce
Tools & Teardowns3 min read

Mac Studio M5 Ultra: 512GB to run big models on one desk

Apple's M5 Ultra Mac Studio tops out at 512GB unified memory and 1.2TB/s bandwidth, from $5,499. What that actually buys you for local model hosting.

Apple announced the M6 and M5 Ultra today, and the number that matters for anyone running models locally is 512GB of unified memory at 1.2TB/s. That configuration lands in the new Mac Studio starting at $5,499, and it moves a whole class of open-weight models from "rent it by the token" to "it lives under the desk." The catch is in the bandwidth, not the capacity.

What actually happened

Per Apple's chip announcement, M5 Ultra carries up to a 36-core CPU (12 super cores, 24 performance cores), up to an 80-core GPU with Neural Accelerators, a 32-core Neural Engine, up to 512GB of unified memory, and 1.2TB/s of memory bandwidth. Apple puts it at up to 1.3x the multithreaded performance of M3 Ultra and up to 4.5x the peak GPU compute for AI.

M6 is the other half of the release: Apple's first 2nm chip, debuting in Mac mini with a 12-core CPU, 12-core GPU, dual 16-core Neural Engine, up to 32GB of unified memory, and 170GB/s of bandwidth. Apple claims roughly a 30 percent increase in peak GPU compute for AI over M5.

The Mac Studio release has the part you actually budget against. M5 Max starts at $2,499 with up to 128GB. M5 Ultra starts at $5,499 with up to 512GB. Six Thunderbolt 5 ports, Wi-Fi 7, and Apple's claim that clustering multiple units over Thunderbolt 5 gives up to 3x faster inference. Pre-orders opened today; machines ship September 22, and the 512GB configuration does not arrive until late October.

Why local model hosting matters for your business

Capacity decides what fits. Bandwidth decides how fast it talks. 512GB means a large open-weight model loads without quantizing it into mush. 1.2TB/s means generation speed will still be a fraction of what an HBM-backed datacenter GPU delivers. For batch work — nightly document processing, catalog enrichment, transcript summarization — that tradeoff is fine. For a chat widget on your storefront, it is not.

The math is a break-even, not a win. $5,499 plus the 512GB upgrade against a metered API bill is a real calculation with a real answer, and the answer depends entirely on your volume and duty cycle. A box that idles 20 hours a day is a worse deal than a per-token invoice. Run the numbers before you run the purchase order.

The privacy argument is the durable one. Client records, contracts, payroll, anything with a data processing agreement attached — the reason to keep it on a machine you own is not cost, it is that no third-party retention policy applies. That reason does not fluctuate with token pricing.

Late October is a real date. If the 512GB config is load-bearing in your plan, you are planning for Q4, not September. Build the pipeline against a rented GPU first so the hardware arrives into working software.

Key takeaways

  • M5 Ultra: up to 36-core CPU, up to 80-core GPU, 32-core Neural Engine, up to 512GB unified memory, 1.2TB/s bandwidth
  • Mac Studio starts at $2,499 (M5 Max, up to 128GB) and $5,499 (M5 Ultra, up to 512GB)
  • Pre-orders today, general availability September 22; the 512GB config slips to late October
  • M6 is Apple's first 2nm chip — 12-core CPU, up to 32GB, 170GB/s — shipping in Mac mini
  • Unified memory capacity says what loads; 1.2TB/s says how fast it generates. Size the workload to the bandwidth
  • Local hosting wins on data control first and on cost only above a certain duty cycle

Before you buy a box, find out whether you need one. We size local-versus-API inference against your actual volume, then build the pipeline so it runs either way. Run the numbers, or tell us what you're trying to keep off someone else's servers.

Sources: Apple Newsroom — M6 and M5 Ultra, Apple Newsroom — Mac Studio.

  • #apple-silicon
  • #local-inference
  • #mac-studio
  • #m5-ultra
  • #hardware
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.