AI News Today
← All episodes
Episode 22 · June 5, 2026 · 07:09

Hermes is now FREE…Here’s how!

NVIDIA Nemotron-3 Ultra Is Free in Hermes Agent (Setup + Why It Matters)

The script introduces NVIDIA Nemotron-3 Ultra, a newly released open-source 550B mixture-of-experts frontier model now available for free for two weeks through News Portal and usable inside Hermes Agent. It explains that the model is designed for long-running agentic work (planning, tool use, failure recovery) and is advertised as five times faster while using 30% fewer tokens on agent tasks by activating only the needed experts. The walkthrough shows how to enable it via Hermes Agent’s new Mission Control dashboard by going to Manage → Models and selecting Nemotron-3 Ultra, then testing it in chat. It highlights potential benefits for Hermes “goal mode,” mentions benchmark comparisons against GLM 5.1, Qwen 3.5, and others, and notes it’s also available on Hugging Face for local hosting.

00:00 Nemotron Free Drop
00:37 Mission Control Setup
00:59 Why It Matters
01:48 MoE Speed Explained
02:49 Three Step Install
03:25 Real World Tests
04:18 Goal Mode Power
05:16 Benchmarks Breakdown
06:07 Wrap Up Recap
06:13 Community Pitch
07:06 Final Goodbye

Full transcript

There's a brand new free model with Hermes Agent. I'm going to show you exactly how to use it, how to set up and how it works. This is a Nemetron Ultra 3, and you can just see that they've announced it. So Nemetron 3 Ultra is available for free with Hermes Agent on the news portal, which means you can start building with it and create some awesome stuff with it.

So let me show you exactly how it works, what it means and whether you should actually use it. So this is the new model announcement and literally Nemetron 3 Ultra just dropped today. So this is a brand new model. You can start using straight away.

And one of the biggest improvements with this is it's five times faster, right? So it's designed for tasks and it's five times faster, which is great. Now, if you want to use this, you can go to your mission control like you see right here. And then inside the model section, you can change it.

So you can see, for example, we've selected news, which is news portal with NVIDIA Nemetron 3 Ultra. So this is free for two weeks. And you can have that as a main model. Now, once you've done that, if you go to the chat, let's just test this out.

And you can see here it's working. So we can start using this straight away. Now, if you're wondering, OK, what is the importance of this? How does it work, et cetera?

So this is a 550 billion parameter frontier model and it's open source and it's designed and built for long running agents. News Research is giving it away free for two weeks on News Portal. You can plug into Hermes. You've got a frontier grade coding and reasoning brain running your AI agents for free.

And here's an example of what we built with it just for fun, really. But you can see here, basically, we asked and created a living galaxy that's been built by Nemetron with Hermes agent. You might be wondering how to perform. So this is a frontier model.

It's designed specifically for AI agents like, for example, Hermes. Most models have really built for chat. Nemetron 3 Ultra is built to work so it can run for hours as an AI agent planning, using tools, recovering from failures, deciding what to do next. Now, it's a giant mixture of experts models of 550 billion total parameters, but it only fires the actual steps that it needs for each.

So it's faster, right? And so it's a mixture of experts model. What that means, essentially, is if you ask it for a task, it will just use certain parts of the model, not the whole 550 billion parameters. It's also five times faster, which makes it a lot quicker.

And obviously with AI agents like you're going back and forth with it, you're asking it to do stuff, et cetera. So Nemetron 3 Ultra is great for that. That also means it uses 30% less tokens on agentic tasks and it's open source, which means that if you had an amazing setup, you could run this locally. But if you don't have an amazing setup, no problem.

You can set up news portal like I've shown you today. And it's a joint release from NVIDIA and News Research, post-trained specifically for agent setups, which is exactly why it slots straight into this stack really nicely. It's an engine built to run AI agents for hours without losing the threat. Now, how do you set it up?

So three simple steps. You sign up at News Portal, which is free. Then you connect it inside your terminal, or I would recommend these days, it's easier just to go to the manage section of your AI agent, go to the dashboard, go to models, and then just change it over there. Then you pick the model, NVIDIA Nemetron 3 Ultra.

You also get a lot of other free models inside News Portal. So if you haven't checked this out, it's a great way to just avoid using and get free models. So step 3.7 flash is also available for free, and then you can use it. You can chat, or you can hand Hermes along tasks.

And it now runs on a 550 billion parameter frontier brain for free for two weeks. We've already put it inside the agent operating system, and it works pretty nicely as you can see. So this is an example of what we created with it. We also tested it on some reasoning tasks, like you can see right here.

And it performed pretty well. So it worked the whole thing. Now, interestingly, here's something that makes it really good. So it thinks in terms of long horizons.

So you can plan, you can do multi-step tasks. And some people say a 550 billion frontier model is too expensive, but this is free. You might say free models are weak models, but actually this is an open frontier model from NVIDIA and it's tuned for AI agents. And you might also say, I don't want to swap the model inside Hermes because it's too much work.

But I've shown you how quick and easy it is to do. You just go to manage, then you go to models and then you change it out. This is because of the new mission control dashboard that just came out from Hermes agent yesterday. So with that, everything is easier in terms of changing it.

You don't need the terminal anymore, which is great as well. The other thing I think this could be very powerful for is the goal mode. So with goal mode, obviously Hermes agent, it works at a task. You give it a big task and a big mission, and then it will have a go at that task for 20 turns before stopping.

And there's a judge that basically analyzes each turn, each attempt from Hermes agent to see if it's actually been completed. If it's not completed, then Hermes agent just keeps going until it finally gets the job done. And the judge judges it's actually completed. And so because this is a big model and because it is a model that works on log horizons, you could give it a task, like for example, build a website or complete my SEO strategy.

And it could just go for 20 turns or 50 turns or whatever you set and complete that for hours without you having to do anything in between, which is pretty powerful. It's also designed to excel at complex tasks like coding and deep research. So long running agents can spend the time planning, using tools, recovering from failures, deciding what to do next. You might say, how does it perform in terms of the benchmarks for this sort of stuff?

So if you look at the benchmarks here, this is Nemetron 3 Ultra versus GLM 5.1, KimiK 2.6 and Kuen 3.5. And you can see in terms of agent productivity benchmarks, it's right at the top there outperforming GLM and Kuen 3.5. In terms of instruction following, it's everything. If you look at long context, it's outperforming everyone as well.

And then also, if you look at long horizon planning, it's outperforming KimiK 2.6 and Kuen 3.5, it's not outperforming GLM 5.1, but that's just something to bear in mind. So it excels at professional work tasks, long context, instruction following and agent productivity. And you can see here, they actually post-trained it for AI agent harnesses. So for example, OpenCore is another model you can plug this into.

It'll be pretty powerful for it. And it's available on Hugging Face as well. If you wanted to host it locally, you can see it right here. So that's basically how to use it, what it is, how to get it for free, how to run Hermes agent for free, how to set it up as well, if you want to get my full agent operating system for Hermes agent, OpenCore and everything else, you can get that inside the AI Profitable Audium, link in the comments description or go to the AI Profitable Audium.com.

This is my AI automation community that's focused on helping you save time, grow and scale with AI automation. Inside the community, you can ask questions, get help and support. I answer the questions personally inside here. Plus there's always people online 24 seven, which means you get help from everyone inside the community.

Inside the classroom, you get access to all of my new daily tutorials and best training. So if you're a complete beginner, you can go from beginner to expert here. If you want new daily updates, you can get out of here. And also if you want my agent operating system for managing all this stuff, you can get it and we update this daily.

Plus you get a zip file you can just quickly install. Inside the calendar, you get four weekly coaching course. You can get on these coaching course, get help and support in real time. We also have a map where you can meet people in your local city using AI automation and AI agents like Hermes and OpenCore too.

And this is all inside the AI Profitable Audium. Hope to see you inside there. Cheers for watching. Bye bye.

More episodes

Browse all episodes →