The script covers a new OpenRouter Fusion API update that runs a prompt across a parallel panel of up to eight models (with web search and bash tools), then uses a judge model to extract consensus, contradictions, unique insights, and missing coverage before returning one fused answer. Fusion is presented as a way to boost benchmark performance and reduce token costs versus relying on a single frontier model, with tests on 100 hard deep-research tasks showing much of the lift coming from synthesis rather than diversity. Examples compare solo models versus panels, including a “budget panel” of cheaper models landing within 1% of Claude Fable 5 on intelligence tests, and demonstrations of using Fusion in chat and via API to generate outputs like SEO research and a clean landing page.
00:00 Fusion Update Overview
00:51 Panels Beat Solo Models
01:43 Budget Panel Near Fable
02:23 How Fusion Works
03:05 Live Panel Demo
03:53 Benchmark Results Breakdown
05:02 API Integration Ideas
05:50 Boardroom SEO Example
06:56 Judge Fusion Output
08:05 Draco Benchmark Explained
09:10 Landing Page Results
10:28 Wrap Up And Offers
Full transcript
There's a brand new update from Opum and Fusion that basically allows you to achieve, according to this, according to this, Fusion allows you to achieve fable level intelligence, but at half the price. So you can see the charts right here, in terms of how it performs on benchmarks, and this is pretty interesting as a model. So you can see here, they're basically Fusion on 100 hard research tasks. And there are panels of models that consistently outperform individual models.
So what this is, is basically like you can create a panel of three different models that work together to get better outputs and also to use less tokens. So you can achieve beyond frontier performance with frontier panels. And this is very different to using one individual model, it's a panel of agents that work together. And you can also have, this is really interesting, this is where it gets fascinating.
So you can have panels of budget models, right, cheaper models that can surpass frontier models, and obviously that's cheaper, right, uses less tokens, uses less powerful APIs. So you can see an example of the tests right here. By testing different combinations of models, they found roughly three quarters of the lift that Fusion provides comes from synthesis and one quarter from diversity. So you can see an example of how they perform right here.
So what this means, essentially, it's very interesting. So you've got Opus 4.8 solo, so Opus 4.8 working alone. Then you have the benchmark score with Opus 4.8 and Opus 4.8 working together, right. Now, if you have Opus 4.8 and 5.5, you get even better results from the working together, but you can use cheaper models.
So you can use, like even free APIs will be quite interesting to test with this. So what's the most interesting out of all of this is that the budget panel is actually comparable with Claude Fable 5 in performance. So if you have like a panel of Gemini 3.5 Flash, Kimi K2.6 and DeepSea V4 Pro fused together as a panel of agents working together, that would be solo 5.5 and solo Opus 4.8. So bear in mind, these are not as powerful, they're cheaper models, but they would be frontier models because they're working together and actually landed within 1% of Fable 5 on the intelligence tests.
So you might be wondering, okay, how does it work? How does, how do you use this? So when you send a prompt to Fusion and you can do this via API, it fans out to a panel of models in parallel, each with web search and bash tools enabled. So what happens after that is a judge model reads every response and extracts consensus points, contradictions, partial coverage, unique insights, anything they might've missed.
And then you can actually use this, for example, inside a chat, like you can see here, and you can choose which models you have working together. Now you can switch between this. You can select, for example, like a budget panel, or you could have a quality panel, or you could have a custom panel, right? And this is really interesting because now you can have multiple agents working together inside a panel.
So let's test this out. We've got the quality section here. We can plug in a prompt like so. And then from here, it's going to start using all three models at the same time with a judge model coming in later.
Really, really interesting stuff. So these are now generating and they're just working together separately. And you can have up to eight models in parallel working together. So you could have like Opus, Gemini, obviously you can't use Fable anymore, that's gone.
But the difference here is like you could achieve potentially Fable level intelligence with one judge that fuses and gives you one answer. And that's the cool thing as well, is like you don't have to check, you know, five different answers at the same time. The judge fuses the models and then gives you one answer back and you can have eight models in parallel working. So we've got that working over here, as you can see, and so you score higher on benchmarks.
And this was interesting as well. So there's a full breakdown on it here in terms of what they found and how it works. And these are the tests that we've done. Well, not we've done, they've done.
So as a fusion, Fable 5 and GPT 5.5, synthesized by Opus 4.8 as a judge, got the highest score on these benchmarks. But if you look at these benchmarks, Opus 4.8, GPT 5.5 and Gemini 3.1 Pro, synthesized by Opus 4.8, scored within 1% of Fable 5 with GPT 5.5. Now, if you look at Claude Fable 5 Solo, that scored 65.3%. And so the fusion without Fable can score higher than Fable 5 alone.
And so if you're like, oh, you know, I wish we still had Fable around, you know, I don't know when it's going to come back, etc. This might be an interesting way to do that. Now, some people watch this and be like, you know, it's benchmarks and I'll try some whatever. I would just try it out for yourself.
See what you get back. See how it goes. I like, I don't know if it's, it's at the point now where it's like so early with this stuff that we don't know. But it is one of the most interesting techniques I've seen to get the most out of these models without having to do anything too special, right?
You can just get the API and it's working together. Now, the interesting thing is you can use this as an API model. So you can use OpenRoute to Fusion and then plug that into whatever you want. So as an example of that, you could get free Claude code, which allows you to plug in the API from Fusion.
And then with Fusion, you could use the power of Claude code's harness to get something really interesting out of this. So we've actually set up the Fusion boardroom inside our agent operating system. And so we can ask a question here and then we can see what we can get back. Now, if you want to indicate, how does that work?
Well, you can see an example of how it created this actual tool where you can paste in an OpenRoute API key, ask it anything, and then it will come back and come back to us. And it was really nicely designed as you can see here. So this is the full tool that we created, where it's kind of like a boardroom tool with Fusion plugged in. So if we go to the boardroom, we can plug in an example for SEO.
So you might say, right, act as a, you know, act as an SEO content council for the keyword, best AI community, search the web for what currently ranks, then for synthesize search intent, the angle competitors are missing, a recommended H2 outline, three questions every article forgets to answer. And then we can click convene and the judge and the panel will start working together to do that via the API with OpenRouter, which is pretty mind blowing when you think about it, because now you've got something that is potentially Fusion level. But at the same time, you don't have the problem of all the, you know, if you're using Fable 5 on an API, like, well, that would be insane for using up tokens, right? And you can see the output here if we go directly inside the chat.
So we've said, you know, create this, blah, blah, blah. And then it's gone off and created the outputs. Now, interestingly, you know, if you did this, Opus came back without the page. But if we look at OpenAI and we look at Google, they've created a full landing page for that particular example.
Now what happens is that we do the analysis. So what's happened here, if we look at this, is that Fusion filters with the judge and it analyzes the agreement, key differences, partial coverage, unique insights, and anything that was potentially missed. So you can see a bunch of examples here. We've got the unique insights, the partial coverage, and each one of these models is working together as a team to give the best outputs.
I mean, if you have more minds and they're all good working together, you know, three brains probably better than one. And also, this is an interesting way to orchestrate all your agents together or all your APIs together in a way that's very, very easy because you've got an API working together. And then what we have from here is a fused answer. So this is a fused answer from the judge.
If we read the full response, it's come back with the HTML code, as we can see. And then it comes back to us. Now, it's actually still writing that out. So it's still processing the HTML.
So it does take a bit of time to get the outputs. But that is a really fascinating, I've never seen AI really used in that way. And it's, it's really cool because you've got an orchestra of APIs just being effortlessly combined into one system. Again, it is a bit slower because it's got that section where it fuses the answers together.
But I think if you want the best possible outputs, that's a great way to approach it. Now you might be wondering, okay, what benchmark did they use for this? They use something called Draco to test reasoning, tool calling, and nurse. So it is something that could tell the difference between a model that sounds for it and one that actually is.
So they use Draco by perplexity, which basically contains 100 deep research tasks, spanning 10 domains. So academic research, technology, UX design, et cetera. And then it checks all of them. And so for example, Opus 4.8 versus itself gets a 6.7% jump in improvement versus the original.
So on the benchmarks, Opus 4.8 as a two model panel scored 65%, 65.5%, which is much higher than Opus 4.8 solo, which is 58.8%. It's a lot of fun, a lot of fun to test and try out. You can see our team working right here. Again, it does take a while for the outputs to come back and you can see that we've got that now.
And then it gives us a breakdown here. So it's like, here's what I'd change, here's the design highlights of the page you wanted to create, here's a note on the stack, et cetera. And then we can test out and that's the page that's been created. And it's pretty nice, like the UI, look at that.
So you created this full page and the UI, I mean, that is one of the most impressive things about this is like, usually, you know, if you're using Opus 4.8, particularly for websites, the design and the UI of the page is not that nice. But if you look at this, like it looks good. It looks really nice, really nice and clean, easy to use, easy to set up. And literally all we did as a prompt to generate something as nice as that was just say, create a beautiful landing page for an SEO agency.
That was it. So definitely worth testing out. Now we've already plugged it into our agent operating system along with everything else, as you can see right here. But basically you've learned how you could potentially, according to the benchmarks, achieve Fable 5 level intelligence with a simple API that uses fusion of all the models and then judges whether it's actually good or not.
And it's kind of similar to how Go works. So if you look at Go mode, what Go mode does with like, for example, Hermes or Claude is the agents get a turn each at achieving the goal. They submit the work, the judge checks if it's good. And so if you look at, for example, how Go mode is working and also how fusion mode is working, it's like something that's really good as best practice is to have a judge in place who checks the work, quality controls it, and then fuses all the answers or tells the agents to iterate from there.
So pretty powerful. So thanks so much for watching. If you want to get my agent operating system, like you see right here for orchestrating agents, we also have paperclip built in. We have a AI agent mastermind, which is quite similar, where it's a group chat between your agents.
We could even plug fusion into there if you wanted. Then also we have the fusion section down here. You can get that inside the AI profit boardroom, which is my AI automation community focused on helping you save time, grow and scale with AI automation. So inside the community, you can ask questions, get help and support whenever you need to.
I personally answer the questions inside there. Plus there's always people online 24 seven. So you can get help whenever you need it. Inside the classroom, you can get access to all of my best and new training.
So if you're a complete beginner, you can go from beginner to expert here. If you want to see the new advanced stuff, we add new daily tutorials like you can see. And then we also have the agent operating system where you can get the video tutorial, the last update date, and then you can also get the zip file to install like so. Inside the calendar, you can jump a week coaching calls, get help and support in real time.
There's four weekly coaching calls a week, which means you can share your screen. You can ask questions. You can meet other cool people doing similar things. Inside the map, you can meet people in your local area using AI agents like you.
And this is all inside the AI Profitable Boardroom. So thanks so much for watching. I'll see you on the next one. Cheers, bye.
More episodes