Sakana's Fugu Ultra v2 Routes Around Astra and Fable
The Tokyo lab shipped two orchestrators on September 11: a cost-first Fugu Max at $2 per million input tokens, and a capability-first Ultra v2 that takes best or joint-best on five of eight benchmarks with Claude Fable 5, Fable 5.1 and GPT-6 Astra deliberately kept out of its pool.
Sakana AI released two models on 11 September, and neither of them is a model in the usual sense. Fugu Max and Fugu Ultra v2 are learned orchestrators: you send a request to one OpenAI-compatible endpoint, and Fugu decides which of the models in the pool behind it should actually do the work, in what order, and how many of them to spend. The Tokyo lab has been selling that idea since it made the first Fugu generally available in June. This week it split the product in two and pointed each half at a different question.
Fugu Max optimises for the cheapest acceptable answer. It lists at $2 per million input tokens and $6 per million output, with cached input at $0.25 and web search or fetch billed at $0.007 a call — output pricing Sakana puts 40% to 60% below Claude Sonnet 5, GPT-5.6 Terra and Kimi K3. Fugu Ultra v2 optimises for the hardest multi-step work and prices accordingly: $5 in, $30 out, $0.50 cached, stepping up to $10, $45 and $1.00 once a request crosses 272,000 tokens of context. Both carry a one-million-token context window and up to 128,000 tokens of output, and subscription tiers from $20 to $200 a month cover every variant.
Sakana's benchmark claims are mostly phrased as counts of wins rather than scores, which is its own kind of tell. Fugu Max takes the best overall result on six benchmarks — Terminal Bench 2.1, GPQA Diamond, AA-LCR, GDP.pdf, AutomationBench and SWEFish — and the company says it expands the cost-performance frontier on seven of ten. Fugu Ultra v2 takes best or joint-best on five of eight (GDP.pdf, Chartography, SWEFish, DeepSWE and Toolathon) and lands top-two on seven of eight. Only two head-to-head figures appear as text anywhere in the release: 74.3 on DeepSWE, and a large lead on chart comprehension.
That Chartography gap — 48.3 against 27.3 for Claude Opus 5 and 29.5 for Claude Fable 5 — is the kind of result orchestration is supposed to produce. Reading a figure well is a narrow skill, and a router that can hand the task to whichever pool member happens to be good at it should beat a generalist that has to be good at everything. It is also, as with every number here, Sakana's own measurement against models it does not control.
The pool is where the strategy actually lives. Sakana says Fugu Ultra v2 reaches those scores without Claude Fable 5, Claude Fable 5.1 or GPT-6 Astra in its agent pool at all — the three most expensive frontier models on the market are simply not in the rotation. What is in there is open-weights and specialist models, including Nvidia's Nemotron family through a partnership. The full routing table is not disclosed, and Sakana is explicit that this is by design. Read one way, that is the entire pitch: a customer who buys Fugu is not exposed to any single vendor's pricing, availability or policy changes. Read the other way, a customer who buys Fugu cannot tell which model handled a given request, which is a harder conversation to have with a compliance team than with a CFO.
The operational details are unglamorous and mostly favourable. The model ID is fugu-ultra-v2.0, and Sakana says moving off an earlier Fugu is a single-line parameter change with no SDK migration. The coordinator's training cutoff moved to 28 August 2026, meaning it has been retuned against the current generation of frontier models rather than the spring line-up. Both models are hosted API only — there are no weights to self-host — and neither is offered in the EU or EEA while Sakana works through GDPR compliance, which rules the product out for a large share of European enterprise buyers on day one. A technical report accompanies the release.
The caveat that matters most is the one the rate card invites. Every figure above is vendor-reported and unreplicated, and a price per million tokens is not a price per task: an orchestrator that routes a job to a small model spends far fewer of its own output tokens than a flagship grinding through the same problem, which is precisely the arbitrage Sakana is selling — and equally, a router that decides a hard problem needs five model calls can cost more than the $5 input rate suggests. The honest comparison is what a full workload costs end to end, and nobody outside Sakana has run one yet.
What is clear is the position. Sakana is not trying to build a frontier model and has stopped pretending the pool needs one. A Japanese lab shipping a product whose selling point is that it deliberately excludes Anthropic's and OpenAI's best work is a bet that the value is migrating from the model to the routing layer above it — and that enough buyers would rather own the decision of which model runs than own a relationship with the lab that built it.
Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.