Skip to content
Rush Commerce
AI & Automation4 min read

Sakana Fugu Max: $2 model routing that skips the frontier

Sakana shipped Fugu Max and Fugu Ultra v2 — orchestration models that route across a pool. Ultra v2 posts its top scores with the frontier tier excluded.

Sakana AI shipped two models today: Fugu Max and Fugu Ultra v2. Neither is a model in the sense you are used to. Both are learned orchestrators — one API endpoint that decides, per request, which other model actually does the work. The detail worth your afternoon is buried in Sakana's own release note: Fugu Ultra v2 posts its headline benchmark scores with Fable 5, Fable 5.1 and GPT-6 Astra excluded from the routing pool. Model routing just produced frontier-class numbers without the frontier tier in the room.

What actually happened

Per Sakana's release post, Fugu Max is priced at $2 per million input tokens and $6 per million output tokens. Sakana reports it taking top scores on six benchmarks — Terminal Bench 2.1, GPQA Diamond, AA-LCR, GDP.pdf, AutomationBench and SWEFish — while integrating open-weights and specialized models, including NVIDIA's Nemotron family, into the pool.

Fugu Ultra v2 is the expensive half of the pair and aims at hard multi-step work. Sakana reports it leading five of eight benchmarks, including 48.3 on Chartography against Opus 5's 27.3 and 74.3 on DeepSWE. Sakana's framing is that the two are "not separate products" but the same orchestration architecture tuned for two missions: cost and peak capability.

Both ship behind an OpenAI-compatible API, and Sakana says moving between them is a single-parameter change.

Two caveats we would not skip. Every number above is vendor-reported on vendor-chosen benchmarks — Sakana grades Sakana. And the routing pool is not a published roster, so you are buying a decision process you cannot inspect.

Why managed model routing matters for your business

The pitch is that you stop choosing models. That is genuinely appealing if you are currently maintaining a hand-tuned switch statement across four providers.

The cost is that routing is the single biggest lever you have on both your token bill and your failure modes, and this hands it to a vendor whose pool changes without telling you. You cannot pin behavior. A prompt that worked in September can land on a different underlying model in November and quietly change shape — and you will debug it blind, because the thing that changed is not in your repo.

So use the signal, not the dependency. The signal is real and it is the headline: a router hit top-tier scores with the frontier models deliberately left out. If that holds on your tasks, your default route does not need the frontier tier either, and you have been overpaying for the safety of a familiar model name.

The way to find out is your own eval set, run against your own traffic. Point it at Fugu Max, at a cheap open-weights model you host, at whatever you run today, and read the quality column before the price column. Keep the OpenAI-compatible shape at your boundary so any of those is a config change. That is the same discipline we described when Moonshot's cheap tier grew up — the vendor names rotate, the boundary is what saves you.

Key takeaways

  • Sakana shipped Fugu Max and Fugu Ultra v2 on September 11, both orchestration models behind one OpenAI-compatible API
  • Fugu Max: $2 per million input tokens, $6 per million output, with open-weights and specialized models in the pool
  • Fugu Ultra v2 reports 48.3 on Chartography vs Opus 5's 27.3, and 74.3 on DeepSWE
  • Ultra v2 excludes Fable 5, Fable 5.1 and GPT-6 Astra from its pool - the frontier tier is not doing the work
  • All benchmark numbers are vendor-reported; the routing pool is not a published roster
  • Outsourcing routing means you cannot pin model behavior across releases
  • Run your own evals against your own traffic before you change a default route

If a router can skip the frontier tier and still clear the bar, so can your stack. We build model boundaries that make a provider swap a config change, and eval harnesses that answer "is the cheap route good enough" with your data instead of someone's leaderboard. Run the numbers on your model spend, or see how we build the routing layer.

Sources: Sakana AI: Introducing Fugu Max and Fugu Ultra v2, Sakana AI: Sakana Fugu.

  • #model-routing
  • #sakana-ai
  • #model-pricing
  • #orchestration
  • #vendor-risk
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.