Hermes Agent has released Mixture of Agents (MoA) Presets, allowing multiple AI models to run in parallel and be aggregated into a single stronger answer, similar to systems like Fusion and Sakana Fugu. The script explains this as a workaround for gated or limited-preview frontier models (e.g., GPT 5.6 and Claude Fable 5) by improving performance through a “panel of experts” approach rather than relying on one model. It introduces Hermes Bench, noting claimed reference MoA scores 8% higher than Opus and 11% higher than GPT, and highlights top performance from an Opus 4.8 + GPT-5.5 aggregator. Viewers are shown how to update Hermes, select MoA presets via terminal and dashboard, configure via desktop app or config.yaml, and switch presets with provider commands.
00:00 Hermes MoA Presets
00:43 Why Models Are Gated
02:09 Update and Enable MoA
04:06 Commands and Config
04:49 Top Hermes Bench Combos
05:04 Panel Beats Genius
05:47 Fusion and Alternatives
06:35 Build Systems Not Models
07:34 Agent OS Workflow
08:42 Old Way vs New Way
09:39 AI Profit Boardroom
10:58 Wrap Up and Links
Full transcript
Today, we have a brand new update from Hermes Agent, and they've released something called Hermes Mixture of Agents preset. So this is a way to combine multiple models that you can access into one, and then you can choose from a powerful variety of mixture of experts and mixture of agents that work together in parallel. It's kind of similar to if you've come across Fusion, which achieved Fable 5 level intelligence by having multiple models work together. We also saw this with Sakana Fugu that came out this week, and that also achieves Fable 5 level intelligence.
And this is a way of basically getting multiple answers from different AI models working together. And the reason that they're doing this is because if you look, for example, at the release of GPT 5.6, which is a limited preview today, we also have the release of a Claude Fable 5 that is gradually coming back, but only for 100 different approved partners. All of these models are getting gated, which means that, for example, if you want to have Fable 5 level intelligence or frontier level stuff, well, it's very difficult to do that with the new releases that are just in preview. So how do you get around that?
Well, what you can do as an alternative is have multiple agent models or multiple models working together to give you better answers. And this is essentially what Hermes Agent are now releasing. And they're also working on a new benchmark, which is Hermes Bench. And you can see on Hermes Bench, they've announced on Hermes Bench are upcoming agentic benchmark, Opus 4.8 and GPT 5.5 reference mixture of agents scores 8% higher than Opus and 11% higher than GPT.
So basically you can improve the performance of your AI agents by using this system instead of relying on just one model at a time. Now, obviously I can imagine that would use more tokens, but at the same time, you can see that mixture of agents will work better together than just having one model working in isolation. Now, at this point, you're probably thinking, okay, how do we get this working for me? So what you can do is if you go into your terminal, you can type in Hermes model, and this is a new setup.
So make sure you're updated. First thing you want to do is just run Hermes update inside your terminal, or you can do that inside your dashboard. If you go to the Hermes manage section, scroll down and click on update Hermes. Once you've done that, then what you want to do is go to Hermes model.
So you type that inside your terminal. And then from here, you'll see an option that says mixture of agents, and these are named presets. So what that means is basically they have preset options for multiple different models working together with mixture of agents. So these are kind of like pre-made formulas where you can have these models working together.
So if we click on enter here, we can then start switching those around. And then once you've done that, if you go to your model section, click on change inside your dashboard, you will see the model settings here, and you can load mixtures of agents. So you can configure this directly. So this is coming soon.
And they've also talked about Fable 5 level stuff as well, right? So they've said, you might be wondering as we did, if it is Fable 5 level. And obviously nobody really has access to Fable 5, but the Hermes benchmarks look promising according to their own words. So it's a pretty cool way to work around this situation with the preview models.
And it's now available inside Hermes agent, so you can get access to it. They're also gonna be releasing a new benchmark for all this stuff too. And it's also provider agnostics. This is not limited to, for example, like News Portal or Open Router, it's provider agnostic, which means you can plug in whatever you want.
Now, if you want the full documentation on that, they've got that inside the Hermes agent News Portal details. So you can select a preset and you can configure it from your terminal provider services. So for example, you can type in model default provider MOA or model review provider MOA. And these are the terminal commands you can use to switch between them.
You can also use these on agent loops as well. And if you want to configure the presets, you can configure that from your dashboard. You can also do this inside the desktop app and you can do Hermes MOA configure or you can go to the config. Now, if we look at what's performing the best on these Hermes bench scores, you can see the Opus aggregator.
So Opus 4.8 plus GPT 5.5 is scoring the best. And then you got Opus 4.8 and GPT 5.5 below that. So if you're wondering, okay, does this actually work in reality? A panel of experts typically beats one genius.
So if you picture, for example, like one brilliant person answering a hard question alone, and then you have a panel of brilliant people, each one writing their own take privately, and then a sharp chair reads all of them and gives you the best combined answer, well, the panel would win every single time. And that is mixture of agents. The reference models are the panel. The aggregator is a chair.
One question goes in, then several models think, and one clean answer comes out better than any single one of them could give. And it's the same, for example, if you're using something like Sakana Fugu or Fusion, let's take a look at Fusion over here. So all of this stuff that you can see here was created with Fusion directly. And when we test it out side by side versus stuff like Opus 4.8, you can check out GoldieBench.
It outperformed pretty much every single model on the lead board. If we actually open up GoldieBench here, you can see that the mixture of experts, mixture of panel, system, Fusion is outperforming pretty much everything else on the leaderboards here by a long way. So it's a really powerful system that actually works. And there's multiple ways.
You don't just have to rely on Hermes to do this. You could use Fusion. You could use Sakana Fugu, whatever you prefer. Now it seems as well, when I've checked this out, that I don't know if there's a limit on the number of agents and presets you can have, but it seems like it's just two models working together inside Hermes agent.
But the main thing you want to focus on here is like stop chasing the model, build the system instead. So everyone else is like waiting on the next model, the next Opus, the next GPT, the thing that'll finally change everything. But if you look at what happened, a mix of today's models beats the best single model that's not available on the market anymore, right? And there's no new release of that.
There's no gated access. It's just a smarter way to use what's already there. It's more efficient. And that's a lesson that mixture of experts or mixture of agents hands you for free.
I keep saying mixture of experts because I'm so used to saying that with local models. Mixture of agents hands you. So the model is not the moat, the system around it is. And the model is a part you can swap.
The system is the thing you own. And the timing on this is huge, right? You've got Fable 5 in preview, you've got GPT 5.6, which isn't confirmed to come out yet, but OpenAI are hoping they can release it to the public. So the winning move isn't waiting, it's squeezing more out of the models that you already have by combining them.
So this is how you can get them working inside your system. I mean, for example, if you look at our setup here, we have Hermes agent and we can use mixture of agents inside this section, but then we could also switch over to Sakana Fugu and start chatting over here with it using the API. And we can see everything that we've built with it directly here as well. And it's the same, for example, with Fusion.
We can chat with it anytime we want. It's just one click away rather than a separate tab. And we've got a system where everything that we've built with this is all ready to go and preview inside one easy system. And that's really the method now.
You don't want to focus on the model, focus on the system instead. It's really a pattern. I've been running this for weeks now. So, you know, mixture of agents, it isn't like a one-off trick.
It's a pattern that I'm seeing recently. I've already built two systems working on the exact same idea, which is a panel of models fused into one answer. So we've got Fusion inside the agent OS as well. We've got Sakana Fugu, and now we can use mixture of agents as well.
Three different systems that we've built into the agent OS to get the most out of this stuff. And you can see how they would work to get the best possible output so you can from your system. And this really changes everything. If you look at the old way, it's like pick one model and hope it's the best.
Hit that model ceiling and stop there. Wait moments for the next release to save you. Beg for access to the gated frontier models. And then when they finally do come out, you know, it can be quite expensive using the APIs on those.
Whereas if you look at the system, you move beyond the ceiling because you can mix several models into one virtual model. You beat the best single model, today with mixture of agents. No waiting, you can squeeze more from what's already available to you. You go beyond the public frontier with no extra access needed, and you hit frontier quality for a fraction of the cost because you've got these agents working together.
Bear in mind, like you could have cheaper models working together and still get better outputs than if one of those models was working alone. And it's just one command. It's just forward slash MOA, and then you can switch around it. So if you want Hermes agent with the panel built in with MOE, mixture of agents, you got Fusion, you got Sakana, all running on the same idea.
We have all of this inside the agent operating system, which is a dashboard that I run pretty much my whole systems on, right? So the full agent OS SIP, Hermes MOA, yeah, a Fusion, Sakana, all wired in, coaching calls, where I build model panels with you live, daily tutorials, so you're stacking models, not just reading about them, and a room of 3,800 operators running this exact stack, plus the prompts, the presets, and a memo map for your city. You can get that inside the AI Profit Boardroom, link in the comments description, or just go to the AIProfitBoardroom.com. This community focused on helping you save time and scale with AI automation.
Inside the community, you can ask questions, and I personally answer them with daily tutorials. Inside the classroom, you can actually get all of my best trainings, including, for example, we've got trainings on Sakana versus Fable versus Fusion. We have a full system on loop engineering to achieve better outputs. And we also have a full training on Fusion versus Fable 5.
And if you want the agent operating system with all of this built in and ready to go, you can see that right here. You can see it's updated daily. You get a full SIP file, plus a 30 day roadmap to implement it. Inside the calendar, you can jump on coaching calls, get help and support in real time.
Inside the map, you can meet people in your local area who are building with AI agents like you. And it's all available inside the AI Profit Boardroom, link in the comments description, or just go to the AIprofitboardroom.com.
More episodes