AI News Today
← All episodes
Episode 97 · August 2, 2026 · 12:49

Sakana Fugu Ultra 1.1 DESTROYS Fable 5?

Sakana Fugu Ultra 1.1: Multi‑Agent “Council of Models” That Beats Fable 5 (Goldy Bench Tests)

The video reviews Sakana Fugu Ultra 1.1, a new Japan-based model that benchmarks above Fable 5 and uses a “council of models” approach: one prompt is routed by an orchestrator to up to three expert models whose outputs are merged into a single build. The creator tests it on Goldy Bench by generating multiple one-shot game builds (e.g., open-world and flight simulator) to evaluate planning, logic, coding, and UI, noting the main drawback is slow generation (about 15–25 minutes) with no back-and-forth. Side-by-side comparisons show Fugu Ultra producing smoother, less buggy, better-looking results than Fable 5. The script also mentions regional blocking in the EU/UK, alternatives like OpenRouter Fusion and Hermes mixture-of-experts, and integrating everything into an Agent OS with saved workspace, plus access options via console.sakana.ai or OpenRouter and a pitch for the AI Profit Boardroom.

Full transcript

Today we're going to be looking at Sakana Fugu Ultra, which is a new model that's just dropped. This is from Japan and if you look at the benchmarks here, Sakana Fugu Ultra 1.1 is actually beating Fable 5 on the benchmarks. Now I've run several tests on it, I'll show you what they look like in a second, but yeah basically Sakana Fugu 1.1 just dropped and it's unlike any other AI model you've used because you can give it one prompt and it sends a whole team of AI experts to build it for you. I actually put that team through the hardest game builds on my benchmark and it actually matched the best model on the planet on its very first try.

So I'm going to walk you through some of the top builds so you can see it yourself. Now if we go over to GoldieBench and we check out the stuff we've created, let's have a look at some of the stuff we have built with this. We've had it running for hours just to see what we can create with it. So this is Fugu Ultra 1.1 and you can see some examples here.

So it creates some pretty nice stuff, I mean this game is super fun to play, 3D looks pretty cool. I'll also compare it to Fable 5 and the same sort of tests in a second so you can compare them side by side. So let's have a look at the next one, this is like an open world game as you can see. You might say, well why don't you use it for business automations?

I usually do, but the difference is that when you're using this for coding out games you get a good feel for how well it can plan, how well it's good at logic, how well it can code etc. And also what the UI is like when it designs stuff. So it's a really good test of what works and what doesn't. This one is pretty cool, like wow, it can create some amazing stuff as you can see.

And bear in mind this is one shot, so what happens with Sukana is that it uses several different AI models to put things together and then create them. So for example if we look at this right here, it's basically got a council of models and that's how it gets better outputs than for example Fable 5 on its own. And if we have a look at this one, it's kind of like a Skyrim style game. Fun to play, easy to use, open world.

All this stuff is really really good. The only downside with Sukana, and this one actually blows me away, is that it's one shot. So you can't go back and forth with it because it takes like 15 minutes to generate one shot. So when we're using this, all this stuff that you're seeing right now was created first time round.

We didn't have to go back and forth here and fly that, but it creates amazing stuff. It's like a parachute game as you can see. Pretty cool. I like the detail and everything like that on the game.

And this is a flight simulator. Pretty nice. Everything worked first time round as well. Let's have a look at this one.

Yeah, yeah, pretty cool. Alright, so how does this work? Basically you give Sukana Fugu a prompt, then you get the orchestrator, that routes to one to three experts, and then you can use the system to get one build. So it merges all of the answers together, and the expert team is working together separately, and then these get fused inside one final answer.

So if we have a look at this system, we've actually got it inside the agent operating system. If we scroll down to Sukana Fugu, we can ask the council of models questions over here, and then we can see everything inside our workspace over here, right? So all this sort of stuff has been built directly using Sukana Fugu, and all the builds you saw before. So everything that we create with this as well, it lands in our workspace so that we can reuse it later.

Now, why do we do this? Well, the thing is, there's a one brain ceiling problem. So every model you use has the same hidden limit. One brain does the whole job, so the same brain plans your physics, paints its graphics, wires its scorebook.

And when the job gets big, that one brain can start dropping things, for example, like context. And so the graphics might land, but the controls break, or the physics might work, but part of it never renders. And so we've all been there. Like a long build comes back 90% right and 10% broken.

Now with Sukana Fugu 1.1, particularly the ultra model, you stop using one brain, and you route the job to a team of experts and merge what comes back. And that's Fugu 1.1. Now you also might say, okay, Fugu and multi-agent sounds a bit like it might be slow. It is a lot slower, right?

So it takes about 25 minutes of orchestrated thinking. But if you need like the best quality outputs, and you need something better than Fable 5, this is well worth looking into. One thing to bear in mind as well is in the EU and the UK, this seems to get blocked. So it doesn't seem to work.

I think that's because UK hasn't allowed it, or sometimes like AI models just don't work in the UK. But if you're in Asia or the US, you can run the request and it actually works, which is pretty cool. And if not, if you want an alternative to this, which isn't using Sukana, you can actually use Fusion. And Fusion is a similar sort of idea.

This comes from OpenRouter. It's API based again, and you can use the console models to get better outputs that way. So everything that we built with it was pretty cool. If we have a look at the benchmarks here, and we compare it against, for example, Fable 5, we've got Claude Fable 5 versus Fugu Ultra 1.1 here.

So this is the game that Fable 5 created, which is pretty nice, but it's a little bit laggy in parts. And it doesn't seem to work that well. When it comes to actually using the controls, you can see it just lags a little bit, right? Whereas if we have a look at this example from Fugu Ultra, it created a better, less linear version with better graphics, it's smoother to run, easier to use, and a lot less buggy, right?

And so because it's using a combination of different models, you tend to get better outputs. You can also use something like mixture of experts directly from Hermes, and that's another way of combining models. If we have a look at the flight simulator, for example, this is Fable 5, which doesn't have that much detail, not that interesting. And then if we compare that to, for example, Sakana Fugu Ultra 1.1, you can see it's a lot more fun to use, it's a lot more fluid, it seems to work better, the controls are nicer, etc.

So it does actually create better outputs side by side versus Fable 5 and any other model that I've tested, actually. By the way, if you want to test and try this stuff yourself, check it out on GoldieBench, you can see a leaderboard of all the models, and you can also go to the compare section and actually play everything that we created. Also, what I like about this system is because we've built it into the AgentOS, we don't have to go and grab the API from somewhere else and then plug it into, for example, Claw to orchestrate it or anything like that. It's all inside this one tab here so that we've got it ready to go with our workspace and everything saved there, along with the chat here where we can use it, right?

So it's just ready to go whenever we need to, which saves a lot of time. Now, some people think, like, I should just pick one model and stick with it, particularly, for example, Fable 5 or something like that. Honestly, every model on GoldieBench wins somewhere and loses somewhere. So the people winning with AI right now run a stack like the AgentOS and then switch and route the jobs to whichever brain or team of brains fits it.

Also, other people say, well, slow models aren't worth it. Like, if Fugle is thinking for 25 minutes, it's not worth it. But if this is just running in the background and you go off and do something else, it actually works really well. And bear in mind, this is beating all the frontier models on the same tests I've run.

And you can see that on GoldieBench. And I've shown you Fable 5 versus Hakana Ultra 1.1 today. Other people say, well, the benchmarks tell me what I need to know. I don't need to test this stuff myself.

But honestly, I never listened to benchmarks. That's why I test out inside the AgentOS. And you can see what we've built with it directly. So the thing that I would say is, like, don't pick your favorite fish.

Own the whole aquarium and route the job, right? When you've got Sakana Fugu, you can route the job to a council of models, to multiple different models and have them all working together, kind of like a mastermind, so that if one makes a mistake, the other one covers up for it. Now, if you're wondering how to get access, you can go to console.sakana.ai, or you can use OpenRouter if it's not available for you on sakana.ai, depending on where you are. And then you can store it inside the AgentOS as well.

You also might say, well, isn't this expensive? But it's like, well, if you want the best model in the world, it's probably, you know, worth it for those one-off jobs where you need something super frontier. Now, you know, like, if you're just coding out a website, you might not use this. But if you're coding out something big and interesting and powerful, then Sakana Fugu can help you in that because you're getting better outputs than something that's frontier alone.

So just to recap here, you understood the whole setup, right? So one prompt, an orchestrator, up to three experts, and then one merged answer with Sakana Fugu. You saw my tests, and I've actually tried to test it and shown you the evidence that actually is pretty good, right? You know the gotchas.

So for example, if you're in certain regions, you would just use the API on OpenRouter instead of going to sakana.ai. And I've shown you the live scorecard as well. And I've shown you how to build it into an agent operating system so that you can use it and save everything you create with it and come back to it later as well. Now, if you want the full Sakana system and all of this system built for you, you could wire it yourself with the steps above that I've shown you, or you can get it inside the agent operating system, right?

And inside the, and you can get our system inside the AI Profitable Boarding, right? So all of this system that you can see right here with the AgentOS, this is all available inside the AI Profitable Boarding. So you can get, for example, a memory system plugged in. You can get all these different orchestration tools.

You can get, we've got a system for Hermes mixture of agents, which is very similar to Sakana Fugu in the way that works. We have Fusion over here, which also works in a similar way. And we have Sakana Fugu with Sakana 1.1 plugged in as well. So if you want to get this full setup, you can get it inside my AI Automation Community, the AI Profitable Boarding, link in the comments description, or just go to the AIProfitboarding.com.

And if you go to the classroom and then go to new daily updates, you can grab the AgentOS system right here with a video tutorial. You can see when it's last updated and you get the zip file to install it. Plus we add new daily tutorials based on what's actually useful and it's just dropped. Now inside the community, you can also ask questions and I personally answer them every day with a video tutorial.

And then inside the calendar, you can jump on weekly coaching calls, share your screen, ask questions, meet other people using the AgentOS. And then inside the map, you can meet people locally near you who are building with AI agents like you. Now some people are going to say, isn't an agent operating system technical to use? So we actually have over 205 pages of testimonials and wins.

So you can see, for example, Andrea posted, thank you, my mission control has been a game changer. I love being ahead of the AI curve again. Thank you, Julian Goldie. And also she set up with tail scale so that she can access the AgentOS anywhere she goes.

Rich said, is anyone else blown away from the AI Profit Boarding? I've been working through the training and I'm genuinely impressed with how everything is organized. So you know, I'm non-technical and all these people are getting great results at the AgentOS. You might not be non, you might be non-technical as well.

It doesn't matter, right? You don't need to be technical to use this. You might also think, okay, well, I don't want to use Sakana because it costs a lot of tokens. Just use it for the big jobs, you know, the super frontier stuff that you need.

And then finally, some people say, well, doesn't an AgentOS use a lot of tokens? But we actually have a full system inside the AI Profit Boarding in the classroom on how to use the agent operating system for free. So if you're on a budget or if you have limited tokens, no problem. You can use this system right here to save tokens and use an agent operating system for free.

So either way you can win with it. So thanks for watching. Hope to see you inside the AI Profit Boarding. See you on the next one.

Cheers.

More episodes

Browse all episodes →