Full transcript
So we have a brand new update from a Japanese AI lab called Sakana, and they've created something called Sakana Fugu. I probably totally mispronounced that, but this is designed to be a full multi-agent orchestration system accessible via a single model API. And this has another model inside it called Fugu Ultra that matches the performance of Fable and Miphos, apparently delivering frontier capability, as you can see right here. Now, basically what this does is it, it's quite similar to Fusion, if you've seen Fusion.
What it should have is basically a mix of models working together, and then they create an answer, right? So instead of just having like one model, like Fable 5, you would ask a question, and then this has a panel, which is a multi-agent panel API that competes head-on, right, and then it synthesizes the answer, and you get one answer. I mean, it literally just dropped an hour ago, so go easy on me, guys. But what I will say is we've already tested out on three different things, and I would always say test out yourself.
I'll explain more on that in a second. So we can see an example of a website we built here. It actually looks super nice. I'll show you how this compares on GoldieBench with everything else that we've created recently, but the website itself looks super nice.
So that was one example. We also have, for example, this maze game, as you can see right here. That turned out super nice, and again, I'll show you how this compares with other models because we always use the same tests and the same prompts with every single model, but it's pretty powerful stuff and actually looks really good, and then we've got this one as well, which is pretty mind-blowing. This is like a simulation of the galaxy, and the quality of the outputs here is super nice, super nice stuff.
So if you're looking for a Fable 5 level, model, or you're looking for Fable 5 level outputs, this could be it, but again, I would test out yourself. I've already tested out myself, as you can see here, and I get the idea of this because we've actually used something else similar recently, and it's called Fusion, right? Now, if you've never used Fusion before, it's a similar sort of idea. You have multiple models as a panel, and then they go for a judge, which fuses it, and then you get one answer out.
Now, if you're wondering, how does Sakana perform on benchmarks? Let's have a look at this. Let's pull this up right here. So we've got Terminal Bench, and you can see Fable 5 scored 80.4 versus 80.2 and 82.1 versus Fugu, and on all the benchmarks here, it's pretty much outperforming Fable 5.
SWE Bench Pro, Fable 5 destroys both Fugu and Fugu Ultra, but on most of the benchmarks, they're either pretty even or it's being outperformed. So you can see, for example, Live Code Bench here, Fugu Ultra, 93.2, 92.9, and 89.8. So pretty impressive model, pretty powerful stuff, really interesting idea. We've already tested it, like I've said before.
Now, if you're wondering, okay, how does this compare against everything else? Let's have a look head on. So let's, for example, compare this versus GLM 5.2, and we can see the same tasks side by side. So this is GLM 5.2, this is Fugu Ultra.
Let's have a look here. So you've already seen this demo, and then let's have a look at the version from GLM 5.2, which is still a little bit buggy, as you can see right here when you're comparing it on benchmarks, and pretty difficult to navigate and move around. Now, if we have a look at the next example here, so this is the website. So GLM 5.2 versus Fugu, Fugu looks super nice.
As you can see, nice animations, nice colors, nice UI, et cetera. And then if we compare that versus GLM 5.2's output, which still looks nice, it's just not quite got that same touch. It just doesn't look quite as nice. So side by side on the outputs here, I would say that Fugu is winning on the benchmarks.
Again, test this stuff yourself, see what you think. Let's have a look at the galaxy example here. So this is the living galaxy, the living spiral galaxy, where we can zoom in, we can zoom out, we can move this around, et cetera. And then if we compare that versus GLM 5.2, looks totally different, right?
Totally different. I would say, which one is more interesting? Which one is more beautiful? I would go with this version right here.
It's just way more interesting to use on a deeper level. And the outputs are pretty amazing, pretty inspiring stuff. I've also got more tests running in the background. So we're testing this out.
Let's compare it versus Opus 4.8 as well. So this is Opus 4.8's output, and this is Fugu's. Like, which one looks a lot more interesting, a lot more powerful? For sure it's this one, right?
More interesting, better design. I will say just even Fable 5 was not that great at UI, but if you compare them side by side, this one looks a lot nicer. Even like the way you can move the mouse and it has these animations over the boxes compared to this, super boring. Now, one thing you need to be careful of, and we saw this with Le Chatant Fat that came out earlier this week.
Be aware of like, you know, companies scoring their own benchmarks. As we saw in this example, this was kind of like a hoax that went viral. And a lot of smart people were tricked by it, right? By Le Chatant Fat scoring and breaking the benchmarks.
This is everything else. So again, this is why I test it. I've created the GoldieBench. This is why I'd recommend you test it out first to see what you think.
But in the side-by-side tests that I've done, it looks great. It's created some nice stuff. Now, if you actually want to use it as an API, what we've actually done over here is build it into our agent operating system. And this is a great thing about having an agent operating system.
Like Sakana just dropped an hour ago. We can already build it into the agent OS, and then we can use it when we need this stuff. So for example, we've got Fusion as well, which is the other alternative to this and runs on a similar sort of API standard. And what you can see, for example, is like if Fable 5 gets taken down, no problem.
You remove it from the system. If, for example, Sakana comes out within one hour, we've already built it into our agent operating system and we've used it, we've tested it. We can just ask it a question here, and then we'll get the high quality outputs that we were showing you a second ago. And so that's a great thing about having a system like this.
It's just so flexible and it can change whenever you want. And if you're wondering how to use Sakana, so you can sign up at sakana.ai, and then also you would get the API from there. Now there's two different APIs. So there's Fugoo Ultra and Fugoo, two different APIs.
And obviously Fugoo Ultra is the more premium version, but obviously a more expensive API as well to use. They also released a technical support so you can find more details on the research and what they did and how it all works, et cetera. But yeah, you can see that it stands shoulder to shoulder with Fable and Miphos. And I've seen that on our own benchmarks as well.
So it's pretty interesting to see. And I think this is the future really, is like using multiple models together to get the best outputs when you need this. I also like the fact that it uses closed and open models together to test it out. And then what it does is it actually manages like the model selection, the delegation and everything for you automatically.
So you don't need to sign up to the individual models. What you do is you just get the API and it handles everything from there. Now you might be wondering, okay, like when should you use each? So Fugu, the basic version is low latency.
So it balances strong performance with low latency, which means like you can get answers quicker and faster. And it fits naturally into tools like Codex for coding. And you could even use this, for example, inside like, you know, customer face and stuff, which is pretty cool. That's one of the problems with Fusion.
When you use Fusion, it's super slow. You know, you might be waiting five or 10 minutes for an answer. Whereas with this, you can build out quickly. That's how we built those outputs I showed you a second ago and tested it within 60 minutes and built it into our system.
Now, when it comes to Fugu Ultra, this is the flagship model and it's tuned for maximum answers, maximum answer quality on hard multi-step problems. So that's really designed for like super deep stuff, like AI research or all sorts of interesting, deeper tasks that require more complexity. Here's another one that we just had come through. So this is an orbit task, which is pretty cool.
It looks super nice. You can change the timescale here as well. So you can speed up, slow it down. And it has a simulated date as well, which is pretty crazy.
And then you can calculate the days per second as you go around inside the solar system, which is absolutely wild. Now, if you compare this to Claude's output, so this is OPUS 4.8, which one do you personally prefer? For me, I think the Fugu's looks a lot nicer, a lot more interesting. Now you also might wonder, okay, how does it perform versus Fusion?
So let's see and have a look at that. So if you look at the response from Fusion, again, I would say this is one of the nicest websites that was created in all of our benchmark tests. But I would genuinely say that again, Fugu looks a lot nicer in the way that it's created. Like if you look at the animations, the way that the UI is designed, it looks super nice.
It's clean, it's interesting. I would prefer the first website. It's marginal in terms of the output difference, but those are the differences. The other cool thing I like about building it in the agent operating system is like, we can save stuff inside our workspace.
So everything that we create as it comes through gets plugged into this system as well, so that we can check it and come back to it later. We don't lose anything. I think, especially if you're using more powerful APIs, like you definitely want to save the outputs and see, okay, what did we create earlier? And then also, you know, as AI models come out and new ones come out every single day, pretty much, then you can see the progress of what you build and how it works, et cetera, which is super fun.
Now, also what's interesting here is the cost per build. So Fusion itself from OpenRouter is a lot more expensive. So if you can land the same prompts, it's 25% of the cost, which is pretty wild. Now, if you're wondering, okay, how does it compare against Fusion?
Well, the endpoint works pretty nicely. The panel is actually denser, which is interesting. Bear in mind, they're both one shot. So for example, if you're using Fusion for tests or Fugu, like you wait for the panels to reply to you.
So you kind of have this weight between them, which means you just use them in one shot. Like you wouldn't go back and forth with them, like a Cloud CLI or something like that. And also what's really good is like Fusion on OpenRouter, you pay per usage, right? It's API only.
So Kana actually has like a flat rate plan. So if you're doing like high volume agent loops, you probably want the subscription and then you'd go for Sakana instead. So thanks for watching. That is the whole setup from Sakana and Fusion.
Seems pretty fun to use, create some awesome stuff. And we've already built it into our agentic OS system, which we'll release an update on for today. We actually update this daily with new updates. As you can see, we've got Fusion built in there too, if you want to compare them side by side.
And this is a agent operating system where you can plug in all your agents together. If something new comes out, like a new model, you can plug it in. We've got Claude, OpenClaude, Hermes all plugged in together. We have Hermes Jarvis, which is a voice activated version of it that we can talk to and also build with in real time.
We have Kimi code ready to go whenever we need to. And everything that we learn and build and automate, we plug into the system and can like a video agent, an SEO agent, a music agent, a Kanban board, whatever you need. So if you want to get that, feel free to get it inside the AR Profit Bottom community, link in the comments description, or go to the ARprofitbottom.com. Inside the community, you can ask questions, get help and support whenever you need to.
And I personally answer these questions with video tutorials every single day. Inside the classroom, you can actually get all of our best trainings. So we have a complete beginner to expert course here, and we have new daily updates inside this section, including the agent OS system with Sakana. It's updated daily.
You get the zip file and a full guide on how to use it. Plus anything new and interesting that comes out, we actually add a new tutorial for it with a video and a full guide step-by-step right here. Inside the calendar, you can jump on weekly coaching calls, get help and support in real time. Inside the map, you can meet people in your local area who are building with AI agents, and this is all inside the AR Profit Bottom.
Feel free to get it, link in the comments description, or go to the ARprofitbottom.com. Thanks for watching.
More episodes