Full transcript
Run Hermes Agent free forever. So Jemma 4 just dropped a brand new update. It is now actually three times faster. Plus there's a bunch of other improvements to it, including tool cooling, which is perfect for using Hermes Agent.
So we've actually already plugged it into our Hermes Agent over here. And we've got a new profile for Jemma 4, and this is a free local model. The great thing about this is, number one, when we're using Hermes Agent, it's private, number two, it's running locally, number three, it is free to run Jemma 4 and you can plug it into Hermes Agent and it can do agentic stuff. So let's have a look at this example.
We actually said forward slash learn, and then we gave it a guide and you can see here that it's actually broken down the guide and understood exactly how to use this skill in the future, right? So we're using the brand new feature, which is forward slash learn from Hermes Agent and it worked and it's very fast to reply, which means that it's better than ever to use Jemma 4 with this setup. So what has changed here? Well, basically Google fixed the agent skill.
So the improvements rolled out across the Jemma 4 family and the tool calling patch is the one that turns Jemma 4 into, you know, a chatbot that you can run locally into an employee, aka an AI agent that you can run locally. That is the unlock and it makes it way faster and consistent and more accurate on tool creation. So this is something I call the self-owned agent, which means you can run Hermes Agent free forever and Google just fixed the one thing that made local AI agents unreliable. So the way that I set, she set this up is we actually created a new agent profile.
If you want to do that, you can just go to manage inside your Hermes Agent dashboard, and then you can scroll down here, click on add profiles. And basically you want to make sure that you add the model from Olama, right, which is Jemma 4. If you're not sure how to get Olama, you just go to olama.com, download it, and then you'll see Jemma 4. And you can run Olama from Jemma 4 there.
Or the other option is you can actually run local models like Jemma 4 with MLX using Huggy Base and LM Studio. So why would we do this? Well, the thing with using agents is that they use up a lot of tokens, but the problem is that that can get expensive. And so if you're using a local model, you can actually have Hermes Agent with a frontier model like GPT 5.6 as the main brain of the operation, and then that can delegate sub-agent tasks to something like Jemma 4, or if you're feeling adventurous, then you can actually have Jemma 4 as the main brain of your operations, you can have Jemma 4 as the main brain behind your AI agents, and then you can build and automate anything for free and privately using this system as well.
And so if you look at this system, basically everything just happens on your computer. You've got Jemma 4 that can run through Olama or LM Studio, and then you've got Hermes, which is the hands. So Jemma 4 is the brain, Olama is the engine room, Hermes is the hands. And you might also say, you know, free local models, not that great.
I wouldn't say they're anywhere near, for example, Fable 5 or anything like that, even if you've got a, you know, a DGX bulk or something like that, but if you want free and if you want local models that don't use up a lot of tokens, then you can use something like Jemma 4. The other way that you can get the most out of this is not just inside Hermes, but for example, we have a local setup inside our agent operating system where we can use Jemma 4 to build a code locally. We can preview what it built and then everything is saved inside a workspace. So we have the preview tab where we can see what we've created.
We've got the build section where we can see what we've built. And then we have the workspace to see what we've created. You might also say, okay, is this actually good for creating stuff? We actually tested it on a bunch of builds, including a website.
The website was by far the most impressive design. I think it's because we actually gave a custom skill to create the web design as well. It looks beautiful, right? It created this really nice website.
As you can see, it looks really clean. It looks super nice. So Jemma 4 is getting better and I've seen a lot of improvements rolling out from it recently, including MOX, which also makes it faster as well. And so the brain is literally one file.
You know, then you've got Olama or you can run it fully locally as well. You can type a message to Hermes and the brain thinks locally. It can use tools, so Jemma 4 can use tool use as well. And then it just works privately and locally for free as well.
Also, the good thing is about Jemma 4 is that there's lots of different sizes. So depending on your setup, if you have a very lightweight setup, then you can use one of the smaller models. So you'll see that ranges from a minimum of 6 gigabytes, which is super small. And that you can see that's on the MOX or you can use the bigger models, for example, like up to 19 gig.
And you can also check it out on Hugging Face as a model and then get the setup from there too. The other really good thing that this is useful for is agent loops. So for example, if you've got big builds, for example, like the forward slash goal mode, using an agent loop with a local model, it can run all day. It could run on your computer, but it doesn't cost you anything to do that.
And whereas, for example, if you're using like the rented brain models, if you're using the API, you know, every message is costing on tokens, heavy agent loops increase the amount of tokens you use very quickly. Your prompts and files travel to the cloud. You can get rate limited at any point. If your internet goes down, you can't use AI.
Whereas for example, with this system, you know, you don't need to worry about any of that. There's no rate limits or anything like that. Everything's just running locally as well. Also, you can just get other models, for example, like Fable 5 to orchestrate this for you, if you don't even want to manage it for yourself.
And this is something that I'm testing with lots of different local models when it comes to Hermes agent. So I call it the self-owned agent. And there's five layers to this, right? So you can download the brain, it could be Gemma 4, it might be something else tomorrow.
Then you've got the engine room. So you can install Olamo or LM studio to run it. And then you've got the hands. So Hermes is the hands and you can set up a new profile with Hermes as I showed you before, run it.
From there, you've got the actual mission control. So having something like the agent OS is great for testing local models because you have the local set up over here. You have the Hermes agent set up over here, and it's all linked into one brain, one memory with all of your prebuilt workflows, which makes it very easy to test and try new things and get the most out of these local models. And you know, your free brain, it can, it can run on anything.
So it could be, for example, drafting content for you. It could be managing and editing files. It could be doing like research and it could be running agent loops or creating content for you as well. You might also say my math core, my computer is not powerful enough for this, but you can use smaller models.
And you might say, well, this will cost me in quality. Actually, if you look at a recent research paper by Kilo Code, they were testing out whether they could achieve Fable 5 level intelligence without Fable 5, and they actually found that it all comes down to the plan. So if you get your AI agents, like your frontier models to create the plan, and then you delegate that to local models, you get very similar outputs because all the decisions, all the forks have been made in advance and therefore the AI just needs to implement it. And that gives you much higher levels of quality than if you were just using the local model alone.
And you also might say, well, I'll set this up later, but the thing is, this is how you get high head, right? The people wiring free local brains into the stack now are the ones who are number one, reducing the amount of tokens they use and number two, moving forward and learn about this stuff. So if you want to get our full system, you can get that inside the AI Profit Boarding, link in the comments description, it goes to the AIprofitboarding.com. We have the local agent set up, as you can see right here.
We have Permise agent that can run on local models too. We have the full mission control dashboard. You can even, for example, use something like paperclip with local models and orchestrate your agents over here. And we even have an agent mastermind.
So Hermes can be orchestrated by all your agents as well. So feel free to get out, link in the comments description or go to the AIprofitboarding.com. This is by AI automation community that's focused on helping you save time, grow and learn with AI automation. Inside the community, you can ask questions, get help and support.
I personally answer these questions with a video tutorial every single day. And you get help and support from the community. He's always online 24 seven inside the classroom. You can get access to all of our best trainings on this stuff.
If you like Hermes, we have loads of great trainings on Hermes agent. So loads of cool stuff on all this sort of stuff. We actually built out an email outreach agent with Hermes and we've got all sorts of cool trainings and video tutorials that I personally test. And then if you want a full agent OS, you can get over here with the full guide set up and the zip file to install it.
Inside the calendar, you can jump a week of coaching calls, get help and support with time. Inside the map, you can meet people in your local area who are building out with Hermes and local models too. So feel free to get that link in the comments, description, or go to the airpufferborn.com.
More episodes