METAL

Sakana AI Ships Two Models That Pick Other Models

Fugu Max routes each task to the leanest model that can solve it and undercuts rival output pricing by 40 to 60 percent. Fugu Ultra v2 posted its best scores with Fable 5.1 and GPT-6 Astra kept out of its pool.

Sakana AI Ships Two Models That Pick Other Models

Image: METAL

Summary

  • Sakana AI released two orchestration models, Fugu Max and Fugu Ultra v2, on September 11. They share one orchestration architecture, with Fugu Max aimed at cost and Fugu Ultra v2 at peak capability.
  • Fugu Max costs 2 dollars per million input tokens and 6 dollars per million output tokens, 40 to 60 percent below Sonnet 5, GPT 5.6 Terra and Kimi K3 on output, and it took the best score on six benchmarks among models in a similar price range.
  • Fugu Ultra v2 scored 48.3 on Chartography and 74.3 on DeepSWE. The company stated in a chart note that Fable 5, Fable 5.1 and GPT-6 Astra are not in its model pool.

Sakana AI has put a price on choosing between models rather than building a bigger one. On September 11 the company released two orchestration models, Fugu Max and Fugu Ultra v2, and said both run on the same orchestration architecture with only their objectives set differently. Fugu Max is built to produce the same answer for the lowest possible cost, and Fugu Ultra v2 is built to push the score as high as it will go on tasks that take many steps.

If orchestration sounds abstract, think of a call centre. The operator who picks up the phone does not solve every problem alone; they read what the caller needs and hand it to whoever covers cards or insurance. Fugu is the model that sits in that operator's chair, reading a request and deciding on the spot which model should handle which part of it. The company wrote that "the most capable AI will never come from a single monolithic model" and that "it will come from intelligent, collective orchestration."

The numbers on the Fugu Max side are attached to its price. According to the company, Fugu Max charges 2 dollars per million input tokens and 6 dollars per million output tokens, putting its output pricing 40 to 60 percent below Sonnet 5, GPT 5.6 Terra and Kimi K3. On benchmarks it took the best score among models in a similar price range on six tests, Terminal Bench 2.1, GPQAD, AA-LCR, GDP.pdf, AutomationBench and SWEFish, and the company said it pushed the cost-performance frontier outward on seven of ten benchmarks.

The way the price came down was by adding models, not cutting them. The company said Fugu Max widens the pool it can orchestrate to take in open-weights and specialised models at a scale it had not reached before, and that a collaboration with NVIDIA brought the Nemotron family into that pool. It described the mechanism this way: by dynamically routing tasks to the leanest model capable of solving them, Fugu Max "delivers frontier-grade results at a fraction of the token spend." The claim is that the more options there are, the more often the cheap answer turns out to be good enough.

Fugu Ultra v2 aims at the opposite end. The company said it built this one for multi-step reasoning, autonomous research and full-stack software development. The gap opened widest on work that means holding charts and structured material in view for a long stretch: on Chartography, which measures chart reading, it scored 48.3 against 27.3 for Opus 5 and 29.5 for Fable 5. On DeepSWE, which measures software repair, the company claims its 74.3 beat models costing three to five times more.

The most striking part of the announcement is the condition the company nailed down itself. A note under the charts states that Fugu Ultra v2's training cutoff is August 28, 2026, and that Fable 5, Fable 5.1 and GPT-6 Astra are not in Fugu Ultra v2's model pool. That means the scores were not borrowed from somebody else's top-tier models, and it also means they were set with the three strongest models of the moment left out. The company said Fugu Ultra v2 took the best or joint-best score on five of eight benchmarks and placed in the top two on seven.

What this company is digging at is the shape of the performance curve. The premise of the announcement is that the contest should be read on two axes, capability and cost, rather than on model size alone, and the company named pushing that frontier outward as the shared objective of both models. METAL has reported that the competition in coding AI is being fought on that same cost-performance frontier, and this release asks whether an orchestration layer, rather than a single model, can push it.

The name Fugu is not new here. METAL has covered the company putting two models named Namazu and Fugu into its own chatbot, and Fugu has been the name of its orchestration line since then. The two releases split that line into a cheap side and a powerful one.

The technical report METAL reviewed sets out how the orchestration layer was built. In the report published on June 19, the company defined Fugu as a model that reads a user request and assembles an agentic scaffold on the fly, and said it used large-scale fine-tuning, evolutionary algorithms and reinforcement learning together. The report takes as its premise that different providers are specialising in different domains, and sets out to combine that specialisation into a single collective intelligence. If the June release showed that an orchestration layer can match closed frontier models on hard benchmarks, the company says this version raises that ceiling.

For anyone using it, the switching cost is low. The company said both models are served through its standard OpenAI-compatible API, that existing Fugu users need a single-line parameter change, and that there is no migration. They can be turned on from the product page or the console. What has not been disclosed is equally clear: the company did not publish pricing for Fugu Ultra v2, and it named no open-weights model in the pool beyond the NVIDIA Nemotron family.

When reading a scorecard from an orchestration model, the place to look is the comparison set. Most of the top scores the company claims carry the qualifier of a similar price range, and they are written as subsets, seven of ten or five of eight. Beating everything and beating everything at the same price are different claims, and this announcement chose the second one. As long as the composition of the model pool behind those scores stays unpublished, reproducing the measurement from outside is not possible.

The reason orchestration lowers the price is plain. There is no need to call the most expensive model to produce the same answer, and the wider the pool, the more room there is to avoid it. Sakana AI built this announcement around infrastructure that does not shake when a vendor locks in, an API is revoked or a service is cut off, and that argument is carried by scores assembled without one company's three strongest models. The basis on which AI gets bought is moving from the name of the model to the cost per unit of work, and orchestration models stand at the front of that shift.

Comments