Today we’re adding Sakana Fugu Ultrato the gatto model picker. It’s the higher-performance of Sakana AI’s two Fugu models, and it does something different from every other entry in the picker: it isn’t one model at all. It’s a learned orchestrator that routes your request across a swappable pool of frontier models and calls itself recursively when the task needs it.
One endpoint, one API call, one line on your invoice — and behind it, a coordinated team of experts. That maps onto how gatto already thinks about work: you don’t hire a single “AI”, you hire a floor. Fugu Ultra brings the same idea to the model layer.
What Fugu Ultra actually is
Sakana describes Fugu as “a multi-agent system that behaves like a single model.” You send a request to one endpoint. Fugu decides whether to solve it directly or assemble a team of expert models — handling model selection, delegation, verification, and synthesis internally so the complexity never reaches your code. The orchestrator is itself a language model, trained on when to delegate, how agents should talk to each other, and how to fold their work into one reliable answer.
Fugu Ultra is tuned for maximum answer quality on hard, multi-step problems. It coordinates a deeper pool of expert agents when accuracy and depth matter most. Early users have pointed it at AI research, paper reproduction, cybersecurity analysis, and literature and patent investigations — the kind of work that doesn’t fit in one prompt.

“For code review, Fugu Ultra is significantly better than GPT-5.5. It gives comprehensive answers and finds the bugs others miss. Where other tools flag about three issues, Fugu surfaced more than twenty.”
— Software engineer, Sakana beta program
Why it belongs on the floor
We’ve talked about the one-person company as a crew of AI employees. The model layer has been the odd one out — you pick one model and it does everything. Fugu Ultra is the first picker entry that’s structured the same way the product is: a coordinator that delegates to specialists. When you hand Fugu Ultra a hard task, it’s doing at the model layer what Gatto does at the workforce layer.
It’s also a hedge. Sakana built Fugu partly because single-vendor dependency is now a real operational risk — recent export controls on frontier models showed that access can shift overnight. Fugu’s pool is swappable. If a provider restricts access, it routes around the disruption. For a floor that runs your business, that resilience matters.
1M-token context window (standard pricing up to 272K, higher tier beyond). Text + image input, text output. Configurable reasoning effort (defaults to xhigh). Tool calling and built-in web search. $5 per 1M input tokens, $30per 1M output, $0.50 per 1M cached input — the standard rate of the underlying models, no fee stacking. Available now in the Studio model picker under OpenRouter.
Benchmarks
Here’s the part that matters. Across Sakana’s reported suite, Fugu Ultra posts 73.7 on SWE-Bench Pro, 82.1 on TerminalBench 2.1, and 95.5on GPQA-Diamond — clearing GPT-5.5 (xhigh), Opus 4.8 (max), and Gemini 3.1 Pro (high) across coding, scientific reasoning, and agentic work. The full panel:

| Benchmark | Fugu Ultra | Opus 4.8 | GPT-5.5 | Gemini 3.1 Pro |
|---|---|---|---|---|
| SWE-Bench Pro | 73.7 | 69.2 | 58.6 | 54.2 |
| TerminalBench 2.1 | 82.1 | 74.6 | 78.2 | 70.3 |
| LiveCodeBench Pro | 90.8 | 84.8 | 88.4 | 82.9 |
| GPQA-Diamond | 95.5 | 92.0 | 93.6 | 94.3 |
| Humanity’s Last Exam | 50.0 | 49.8 | 41.4 | 44.4 |
| MRCRv2 (long-context recall) | 93.6 | 87.9 | 94.8 | 84.9 |
It’s not a clean sweep, and that’s worth saying plainly. Anthropic’s Fable 5 still leads SWE-Bench Pro (80.0) and Humanity’s Last Exam (53.3); GPT-5.5 edges it on long-context recall; Opus 4.8 is a hair ahead on the CTI-REALM security benchmark (69.6 vs 69.4). And running a crew instead of a single model has a cost — early testers report hard turns can take minutes while the conductor delegates and verifies. But the pattern is consistent: Fugu Ultra’s edge shows up on the messy, long-running, multi-step tasks that don’t fit one model call — exactly the work your employees do in background runs and team missions.
From the beta
- Code review.“Where other tools flag about three issues, Fugu surfaced more than twenty.”
- Security assessment.One scoped instruction drove recon, XSS/SQLi checks, auth review, and a clean report with evidence and retest steps — staying in scope, no destructive actions.
- Persona stability.Strong identity retention across long sessions where other models drift — a property that matters more than raw benchmark scores for agent products.
Try it
Open Studio, click the model pill, and pick Fugu Ultrafrom the OpenRouter section. It runs at xhigh reasoning by default, so it’s best reserved for the hard turns — research, planning, code review, anything multi-step where you’d otherwise swap models yourself. For everyday chat, keep using Gatto 1 or your usual picker default.
You’ll also see a pinned note in your inboxpointing back here. Dismiss it once you’ve tried it — it won’t come back.
