Skip to content
Rush Commerce
AI & Automation3 min read

IBM's $240M Together AI deal: cheap inference lands in 2027

IBM and Together AI signed a $240M multi-year deal for an open-model inference cluster on IBM Cloud, available Q1 2027. Why your token budget can't wait for it.

The cheap open-model inference you're planning to route your workloads to is, right now, a purchase order for chips that don't exist yet. IBM and Together AI signed a multi-year, $240 million agreement today to stand up a dedicated open-source AI inference cluster on IBM Cloud. It comes online in Q1 2027. That date is the story.

What actually happened

Per IBM's announcement, the cluster runs NVIDIA HGX B300 systems with Spectrum-X Ethernet networking, sited on IBM Cloud, with availability in the first quarter of 2027. Reuters reports the initial build is roughly 2,000 Blackwell B300 chips, located in the US. Together AI will use it to serve inference on open-weight models — it currently moves about 400 trillion tokens a month.

Together's CEO Vipul Ved Prakash framed the pitch as "the performance of the best frontier models without the closed-model price tag," which only works "if the infrastructure underneath is fast and reliable at scale." Two details are more useful than the quote. First: Together AI raised $800 million at $8.3 billion five weeks ago and is still renting rather than building — even a well-funded neocloud buys capacity from a bigger balance sheet. Second, per The Register's reporting, Together's chief revenue officer expects the cluster to be sold out two to three months before it enters service. Capacity is spoken for before the racks are powered.

Why it matters for your business

Two practical consequences, neither of them about IBM.

Don't build a 2026 budget on 2027 prices. We keep seeing cost models that assume open-model inference gets steadily cheaper because the announcements say more capacity is coming. More capacity is coming — in Q1 2027, pre-sold, at whatever the market clears at then. Your token bill between now and then is set by capacity that already exists. Price your automation on today's rates and treat any decrease as upside.

Your "independent" vendor has a supply chain too. Together AI is one of the better answers to closed-model lock-in, and it depends on IBM Cloud, Nvidia's shipping schedule, and a contract. That's not a knock on them — it's the shape of the whole market. It just means vendor diversification has to be real: two providers you've actually tested against, behind one interface in your code, with an eval set you own so you can prove the swap doesn't degrade output. A second vendor you've never run a request through is not a fallback.

Key takeaways

  • IBM and Together AI signed a $240M multi-year deal for a dedicated open-model inference cluster on IBM Cloud, available Q1 2027
  • The initial build is ~2,000 Nvidia Blackwell B300 chips with Spectrum-X networking, in the US; Together serves ~400T tokens monthly today
  • Together expects the cluster to sell out two to three months before service — new capacity is committed before it exists
  • Budget token spend on today's rates, and make your second inference provider one you've actually tested, not one you've bookmarked

Is your AI cost model built on capacity that hasn't shipped? We build automation where model calls sit behind one interface, with evals that prove a provider swap holds up before you need it. Run the numbers on today's rates or have us stress-test your stack.

Sources: IBM Newsroom, Reuters, The Register.

  • #together-ai
  • #ibm-cloud
  • #ai-inference
  • #open-models
  • #capacity-planning
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.