Skip to content
Rush Commerce
AI & Automation3 min read

Meta's Arke chip: who owns the silicon sets your token price

Meta puts its MTIA 450 Arke inference chip in data centers in H1 2027 and claims better perf-per-dollar than Nvidia. What custom silicon means for what you pay per token.

The cheapest way to lower the price of inference is to stop paying someone else's margin on the chip that runs it. Meta just put a date on that plan: its third-generation MTIA inference chip, internally called Arke, goes into production data centers in the first half of 2027, with a fourth-generation part behind it. If you buy AI by the token rather than by the rack, the question of who owns the silicon under your model call is quietly becoming the question of what you pay.

What actually happened

Bloomberg reported on September 15 that Meta will deploy MTIA 450 — codename Arke — in its data centers in H1 2027. Broadcom handles design, TSMC handles manufacturing. Meta's VP of engineering Yee Jiun Song told Bloomberg that twelve prototype parts arrived from TSMC on September 1 and landed within 2% to 3% of Meta's simulations, which is a good result for first silicon.

Meta's claim is perf-per-watt and perf-per-dollar, not raw peak: the company says the chips run AI models more efficiently than current Nvidia offerings on those two axes. Treat that as a vendor claim until someone outside Meta benchmarks it. The scale commitment is more concrete: more than a gigawatt of custom-chip compute inside a 12-month window. A fourth-generation part, MTIA 500 / Astrid, finishes design in roughly a month and targets end of 2027.

Note what Arke is not. It is an inference part, aimed at the workloads behind Facebook, Instagram, WhatsApp and Meta AI. Nobody is claiming it displaces Nvidia for frontier training.

Why it matters for your business

You will never buy an Arke. You will buy the consequence of it.

Inference is the cost line that actually scales with your customers — every support reply, every product description, every agent run. That cost is set by three things: model efficiency, utilization, and the margin stack between the sand and your API key. Meta attacking the third one for its own workloads is the same pattern as Google's TPUs and Amazon's Trainium. Each hyperscaler that removes a middleman for itself changes the price it can quote everyone else.

The second-order effect matters more for a small shop. When Meta runs its consumer AI on chips it designed, Meta's open-weight models get cheap to serve on Meta's own hardware and stay ordinary to serve on everyone else's. Model portability starts to mean less than it looks like on paper — the weights are portable, the cost curve underneath them is not.

What you do about it is boring and it is the same answer every time this story runs: do not hard-code a provider. Put model access behind an interface you own — one module, one config value, swappable inputs — so a 2027 price move is a deploy, not a rebuild. And stop signing multi-year committed-spend deals that assume today's rate is the floor. Four vendors are currently spending billions to make it not the floor.

Key takeaways

  • Meta's MTIA 450 (Arke) inference chip enters data centers in H1 2027; Broadcom designs it, TSMC fabricates it
  • Twelve prototype parts arrived from TSMC on September 1 and performed within 2–3% of Meta's simulations
  • Meta claims better performance per watt and per dollar than current Nvidia parts — a vendor claim, not an independent benchmark
  • Commitment is over 1 gigawatt of custom-chip compute in a 12-month period; MTIA 500 (Astrid) targets end of 2027
  • Arke is inference-only — it is not a training replacement for Nvidia
  • Every hyperscaler that owns its silicon changes the token price it can quote you; keep your model layer swappable and avoid long committed-spend lock-ins

Your AI features should not know which chip they run on. We build the model layer as a seam you own, so a provider repricing in 2027 is a config change instead of a quarter of rework. See how we architect it, or tell us what your stack calls today.

Sources: Bloomberg, Investing.com.

  • #ai-infrastructure
  • #inference-costs
  • #meta
  • #ai-pricing
  • #vendor-risk
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.