Japan’s Sakana AI has launched Fugu, a multi-agent system that beats every frontier model you can actually buy on most benchmarks. The catch is in the fine print: on the head-to-head numbers, it still loses to Anthropic’s Fable 5, the very model export controls pulled off the market. Here’s what the data really shows.
Key Takeaways
- Fugu Ultra beats Opus 4.8, GPT-5.5, and Gemini 3.1 Pro on most tests
- It loses to Anthropic’s Fable 5 on every direct comparison
- Fable 5 was suspended under a US export-control order
- Fugu is an orchestrator, not a single model, so scores aren’t like-for-like
- All the benchmark figures come from Sakana’s own testing
What Sakana Actually Launched
Fugu isn’t another chatbot. Sakana AI, a Tokyo-based lab, built it as a multi-agent orchestration system that dynamically coordinates a pool of frontier models rather than running as one standalone model.
The design goal is resilience. By routing each query across multiple models from different providers, Fugu builds native redundancy into the AI stack, so if one model becomes unavailable, the system keeps running on the others.
The timing is pointed. Sakana positioned Fugu as a frontier alternative to Anthropic’s suspended models, leaning hard on a no-lock-in pitch aimed at teams nervous about losing access to a single provider.
It comes in two flavors. There’s a balanced, faster variant simply called Fugu, and a maximum-quality version, Fugu Ultra, that taps a deeper model pool at the cost of speed.
The Benchmarks It Wins
Against the models teams can actually access today, Fugu Ultra’s numbers are genuinely strong. Sakana’s testing puts it ahead of the generally available frontier field on the majority of benchmarks.
The coding results stand out. On the demanding SWE-Bench Pro software-engineering test, Fugu Ultra scored 73.7, ahead of Claude Opus 4.8’s 69.2 and GPT-5.5’s 58.6, a comfortable margin over both.
The wins pile up elsewhere too. Fugu Ultra leads on GPQA-Diamond (95.5), LiveCodeBench (93.2), and TerminalBench 2.1 (82.1), and by Sakana’s tally it tops Opus 4.8, Gemini 3.1 Pro, and GPT-5.5 on 10 of 11 benchmarks. The lone exception is the MRCRv2 long-context-recall test, where GPT-5.5 edges ahead.
For agentic work, that profile matters. Coding, reasoning, and multi-step tasks are exactly where an orchestrated pool can shine, and those are the numbers most likely to catch a CTO’s eye.
Where the “Beats Fable” Claim Falls Apart
Now the fine print. Despite headlines suggesting otherwise, Sakana never actually claims to beat Fable 5. Its own framing is that Fugu stands shoulder-to-shoulder with Anthropic’s top-tier models.
The direct numbers back that caution. Where side-by-side comparisons exist, Fable 5 leads Fugu Ultra by roughly 6 to 9 points: SWE-Bench Pro (80.0 vs 73.7) and Humanity’s Last Exam (53.3 vs 50.0) both go to Fable 5, not Fugu.
The reason Fugu can’t beat it is structural. Fable 5 isn’t in Fugu’s model pool at all, because it was pulled from the market, so the orchestrator is approximating the absence of that model rather than outscoring it.
The honest read is a fine distinction. As one analysis put it, Fugu Ultra is excellent, but it isn’t beating the model it’s measured against so much as filling the gap left by its absence. That’s a very different claim from topping Fable 5 outright.
Two Big Caveats Before You Trust the Numbers
Beyond the Fable comparison, two things should temper how much weight anyone puts on these scores.
First, the source. Every benchmark figure here is Sakana-reported, not independently verified, so they read as vendor-published evidence rather than neutral proof, and warrant your own testing before you rely on them.
Second, the category mismatch. These are orchestration-system scores, not single-model scores, which means comparing Fugu’s numbers to a lone model like Opus 4.8 isn’t a clean like-for-like. An orchestrator coordinating several models is a fundamentally different thing from one model answering alone.
There’s a practical trade-off too. Fugu’s routing is hidden per query, so teams don’t see which model answered, and the fan-out approach carries added cost and latency, factors that can be dealbreakers for compliance, finance, or reproducibility-focused workflows.
Why Fable 5 Isn’t in the Race
The elephant in the room is why the strongest model on these charts is one nobody can use. The answer is regulatory.
The backstory is dramatic. The US government suspended public access to Claude Fable 5 just days after Anthropic launched it as its most capable model, and Anthropic subsequently removed the model from global usage. The company has said the suspension stemmed from US export controls, which the Commerce Department later cleared.
That vacuum is Sakana’s entire opening. With Fable 5 and Mythos 5 off the table for many users, an orchestrator that can’t be geofenced out of existence, and as a Japanese company sits outside the US export directive, has a real pitch to make.
The value proposition writes itself. For teams that were building on Anthropic’s top tier and suddenly lost access, a system offering frontier-adjacent performance with built-in failover is worth a serious look, even if it isn’t quite the fastest gun on paper.
The Bottom Line
Strip away the headline hype and Fugu Ultra is a legitimately impressive launch. It beats every model you can currently buy on most of Sakana’s benchmarks, and its orchestration approach turns provider risk into a feature rather than a liability.
But the “beats Fable 5” framing doesn’t survive the data. On direct comparisons, Fable 5 still wins, and the only reason it’s not in the race is that regulators pulled it. The accurate story is smaller but truer: Fugu Ultra is the best accessible option right now, and it’s betting that availability beats a benchmark crown you can’t actually use. As always, pilot it on your own tasks before trusting any single number.
Digital Trendings is your trusted source for AI news and updates, stay tuned for more.







