gattoBeta
New in the floor

Fugu Ultra, one endpoint, a whole crew.

Fugu Ultra — the red Sakana Fugu fish mark on a dawn glow

Today we’re adding Sakana Fugu Ultrato the gatto model picker. It’s the higher-performance of Sakana AI’s two Fugu models, and it does something different from every other entry in the picker: it isn’t one model at all. It’s a learned orchestrator that routes your request across a swappable pool of frontier models and calls itself recursively when the task needs it.

One endpoint, one API call, one line on your invoice — and behind it, a coordinated team of experts. That maps onto how gatto already thinks about work: you don’t hire a single “AI”, you hire a floor. Fugu Ultra brings the same idea to the model layer.

What Fugu Ultra actually is

Sakana describes Fugu as “a multi-agent system that behaves like a single model.” You send a request to one endpoint. Fugu decides whether to solve it directly or assemble a team of expert models — handling model selection, delegation, verification, and synthesis internally so the complexity never reaches your code. The orchestrator is itself a language model, trained on when to delegate, how agents should talk to each other, and how to fold their work into one reliable answer.

Fugu Ultra is tuned for maximum answer quality on hard, multi-step problems. It coordinates a deeper pool of expert agents when accuracy and depth matter most. Early users have pointed it at AI research, paper reproduction, cybersecurity analysis, and literature and patent investigations — the kind of work that doesn’t fit in one prompt.

Sakana Fugu routing a request across an LLM pool of closed and open models, including Fugu itself
A trained conductor model picks the experts, assigns roles (Thinker, Worker, Verifier), and folds their work into one answer — drawing from a swappable pool of closed and open models. Diagram: Sakana AI.
“For code review, Fugu Ultra is significantly better than GPT-5.5. It gives comprehensive answers and finds the bugs others miss. Where other tools flag about three issues, Fugu surfaced more than twenty.”
— Software engineer, Sakana beta program

Why it belongs on the floor

We’ve talked about the one-person company as a crew of AI employees. The model layer has been the odd one out — you pick one model and it does everything. Fugu Ultra is the first picker entry that’s structured the same way the product is: a coordinator that delegates to specialists. When you hand Fugu Ultra a hard task, it’s doing at the model layer what Gatto does at the workforce layer.

It’s also a hedge. Sakana built Fugu partly because single-vendor dependency is now a real operational risk — recent export controls on frontier models showed that access can shift overnight. Fugu’s pool is swappable. If a provider restricts access, it routes around the disruption. For a floor that runs your business, that resilience matters.

Spec sheet

1M-token context window (standard pricing up to 272K, higher tier beyond). Text + image input, text output. Configurable reasoning effort (defaults to xhigh). Tool calling and built-in web search. $5 per 1M input tokens, $30per 1M output, $0.50 per 1M cached input — the standard rate of the underlying models, no fee stacking. Available now in the Studio model picker under OpenRouter.

Benchmarks

Here’s the part that matters. Across Sakana’s reported suite, Fugu Ultra posts 73.7 on SWE-Bench Pro, 82.1 on TerminalBench 2.1, and 95.5on GPQA-Diamond — clearing GPT-5.5 (xhigh), Opus 4.8 (max), and Gemini 3.1 Pro (high) across coding, scientific reasoning, and agentic work. The full panel:

Bar charts comparing Fugu Ultra and Fugu against Fable 5, Gemini 3.1 Pro, GPT 5.5, Opus 4.8 and Mythos Preview across TerminalBench, CharXiv, GPQA-D, LiveCodeBench, SciCode, SWEBench Pro, Humanity's Last Exam and CTI-REALM
Fugu Ultra (dark red) and Fugu (red) against frontier baselines. Vendor-reported scores from Sakana AI’s technical report — higher is better.
BenchmarkFugu UltraOpus 4.8GPT-5.5Gemini 3.1 Pro
SWE-Bench Pro73.769.258.654.2
TerminalBench 2.182.174.678.270.3
LiveCodeBench Pro90.884.888.482.9
GPQA-Diamond95.592.093.694.3
Humanity’s Last Exam50.049.841.444.4
MRCRv2 (long-context recall)93.687.994.884.9
Vendor-reported, pending independent verification. Bold = best in row.

It’s not a clean sweep, and that’s worth saying plainly. Anthropic’s Fable 5 still leads SWE-Bench Pro (80.0) and Humanity’s Last Exam (53.3); GPT-5.5 edges it on long-context recall; Opus 4.8 is a hair ahead on the CTI-REALM security benchmark (69.6 vs 69.4). And running a crew instead of a single model has a cost — early testers report hard turns can take minutes while the conductor delegates and verifies. But the pattern is consistent: Fugu Ultra’s edge shows up on the messy, long-running, multi-step tasks that don’t fit one model call — exactly the work your employees do in background runs and team missions.

From the beta

  • Code review.“Where other tools flag about three issues, Fugu surfaced more than twenty.”
  • Security assessment.One scoped instruction drove recon, XSS/SQLi checks, auth review, and a clean report with evidence and retest steps — staying in scope, no destructive actions.
  • Persona stability.Strong identity retention across long sessions where other models drift — a property that matters more than raw benchmark scores for agent products.

Try it

Open Studio, click the model pill, and pick Fugu Ultrafrom the OpenRouter section. It runs at xhigh reasoning by default, so it’s best reserved for the hard turns — research, planning, code review, anything multi-step where you’d otherwise swap models yourself. For everyday chat, keep using Gatto 1 or your usual picker default.

You’ll also see a pinned note in your inboxpointing back here. Dismiss it once you’ve tried it — it won’t come back.