AI News Today
← All episodes
Episode 20 · June 4, 2026 · 09:48

Gemma 4: Run Hermes Free Forever!

Run Gemma 4 Locally in Hermes Agent (Free AI Automation + New Web UI Setup)

The script explains how to combine Google’s newly released Gemma 4 12B model with Hermes Agent to run a free, local AI agent designed for agentic reasoning and automation. It shows examples built with Hermes Agent and an agentic operating system, including games, a Pomodoro timer, a color palette, and animations. Setup is demonstrated via Ollama (download Ollama, run the command to launch Hermes with Gemma 4) or through Hermes Agent’s new web UI to select Gemma 4 as the model. It also suggests configuring a stronger main model with Gemma 4 as a sub-agent to save tokens, and notes an alternative free API option via OpenRouter for Gemma 4 26B/31B if local hardware is limited.

00:00 Free Agents With Gemma
00:36 What You Can Build
00:54 Local Setup With Ollama
01:15 Web UI Model Switching
01:32 Main Model And Subagent
02:07 Free API Option
03:01 Body And Brain Explained
04:09 Automation Use Cases
05:13 Offline And Token Savings
06:15 Chat Versus Agent
06:34 Hermes Dashboard Tour
07:20 SOP Pairing Steps
07:53 Benchmarks And Ecosystem
08:24 Recap And Offline Travel
08:51 Training And Community Pitch
09:46 Final Thanks

Full transcript

Gemma 4 and Hermes Agent, you can automate anything for free now. So you can plug Gemma 412B that just dropped today directly into Hermes Agent. Now, why would you do that? Because with Hermes Agent, this is an AI agent and you can plug in the brain like Gemma 4 and then it's free, it's local and it's designed for agentic reasoning.

So this is a powerful way to combine the best of both worlds. A free model with one of the most powerful AI agents in the world. Now, if you're wondering how do you get this set up, etc. I'll show you exactly how to build with it and what we've created using this, plus some examples of agentic tasks.

If you're wondering, okay, what can you build with Hermes Agent? Here's a bunch of examples. We actually built out a game, quite a few games. We created like a Pomodoro timer here, a color palette, even like these animations, which is pretty cool.

So let me show you these. These were all built with our agentic operating system and Hermes Agent combined, like you can see. Now, if you want to get this set up, basically what you can do here is you can go into and Olama have the latest model. So all you do is you download, then you're going to run it.

And then you can run this command, which is Olama launch Hermes model Gemma 4 and get started with this. And you can see this model is ready to go right there. So that's how you can plug it into Hermes Agent pretty quickly. Also, if you're wondering, so basically there's a new web UI with Hermes Agent that makes this even easier.

So if you don't want to use this terminal, you can just go to models over here and then you can change the main model and you can select Olama with Gemma 4 once you've got it installed and downloaded, which is pretty cool as well. So you can configure this. Now, the way that I would actually configure this, and this is using the new web UI from Hermes Agent, which makes it even more easy and powerful to manage, is with this setup, you could have your main model as, for example, like step 3.7 flash, and then you can actually set up Gemma as the sub-agent, right? Because Gemma 4 is not like a super high advanced reason model, but it is fast, it is free, it's local.

And so you would save on tokens if you had it set up as a sub-agent. And then Hermes Agent could have a main model, which is like the agentic reason model, and then it could use Gemma 4 as the local model for sub-agent tasks, the smaller stuff basically, for auxiliary tasks, which is pretty cool. Now, for example, we can also use something called Gemma 4 26b. So if you don't have an amazing setup and you can't run Gemma 4 locally with Hermes Agent, then what you can do instead, and this is pretty cool, is you can get a free API.

And this is Gemma 4 with 26b that you can get for free. And you can also get 30 as well. So you can grab the API for this, and then you can plug it into Hermes Agent. So for example, if we go into the agentic OS, and then we configure an auxiliary model or a main model, we could go down the model list, go to open router, and then we can select Gemma 4 from the list right there.

And you can see, for example, we've got it ready to go. So we can select that as the main model, and now Gemma 4 is ready to go on that, right? And we've also got 31b here too. Both of those are free APIs that you can use.

So if you don't have a local setup, no problem. You can use a free API with Gemma 4 too. Now, if you want to indicate what can you build with this, how can you use it, et cetera, let me guide you through that. So basically, the way that I would look at it is they're two halves, right, on their own.

So Hermes is the body. It's always on, it's living on your machine. It's hands-on, it's working 24-7. Gemma 4 is the brain.

So Google's brand new model that runs for free. And you bolt them together and you get an AI agent that quietly could run your morning brief or clear your inbox or write into your own notes or research for you, even build working apps. Everything below is some examples of what we've created using this, right? And so it runs 24-7 with Hermes.

You've got 12b, which is the new Gemma 4 model. And this is free to use if you're running it locally. So the difference here as well, when it comes to Gemma 4 is unlike a chat that just talks, Hermes is an AI agent that acts, right? So it can implement stuff.

It can manage files. It can run background tasks. It can do browser use. It can autonomously automate stuff on a schedule.

Now, a body needs a brain. So Hermes can do this stuff, right? Hermes can work and implement stuff, but you need a brain and that's what Gemma 4 is. And so Gemma 4 is Google's new open model.

12b is a new model that just dropped today. It's near the quality of models twice its size. And so effectively it's free to run all day. And so it's a free brain that never clocks off because you've plugged it into Hermes agent.

And that's a benefit of bringing them all together. And if you're wondering, okay, what can they build together? So these are some examples. So we could do like a 7 a.m.

morning brief. We could say, okay, every day, brief me on everything that happened. And then it just gives us the top priorities. What happened overnight?

What it's going to do today? Here's another example. So you could actually triage your inboxes, look at your emails and that sort of thing. And here's another one.

We could actually do a weekly review. So we could look at everything inside my memory, which we've actually got plugged into the agent operating system as well. So if we go inside the memory section here, we have used Obsidian as a second brain for Hermes agent and all our other AI agents. So this is plugged into Claw and into Hermes.

And then what we can actually do is get Hermes to review that on a schedule using Gemma 4 as well, which is pretty cool. You could also use it for research too. So you could point it at a topic and then it just goes off and researches stuff. It could come up with content ideas for you too.

And it can build things, right? So for example, it created this pretty cool animation right here just for fun really, but it's interesting to see what it can do and how it works. Here's another one. So this was like your classic sort of snake game test, but it actually came out pretty nice, which is cool as well.

Obviously one thing to note here is like, Clawed Opus 4.8 is going to be way more powerful than something like Gemma 4, but it's cool to see like a small lightweight local model that runs for free. And also you can use it offline. You can't use Clawed Opus offline, but you could use, for example, Gemma 4 offline. So that's a big difference as well.

And so some people want to say a capable AI agent, it uses up a lot of tokens, but if you've got Gemma 4 as a brain, and bear in mind, you can run it locally or you can get the free open router, no problem, right? Because with Gemma 4 as a brain, a full day of AI agents is free. Other people say AI just gives me texts, chat GPT, it just goes back and forth. It doesn't really do real work.

Actually, if you plug a brain like Gemma 4 into Hermes agent, then it can do everything. It can look at your email, so you can do computer use. It can open up apps locally and then build stuff for you. It can use your second brain memory, right?

That's the difference between having a chat and having Hermes agent. Now you can use Hermes agent for free locally too. And then other people say my notes and my files, they go to some big data center in cloud, but actually this is a local model that you can use too. Now here's the old way versus the new way.

So a normal sort of chat is like you do the doing, right? So if you're using chat GPT, for example, you open it, it waits, hands you the words to copy and paste, forgets you when you close it. And also it runs on a subscription and that sort of thing. With Hermes and Gemma 4, it does the doing for you.

So it runs on its own, on a schedule. You can actually manage this. For example, if we go back to the agent OS and then we go to manage and then we go to schedule tasks here, we can create new schedule tasks and we can manage everything inside one place. This is the new Hermes dashboard that just dropped today.

And so you can manage everything in one place, all your skills, all your scheduled tasks, all your logs, all your sessions. You can switch models if you need to switch something else. And if you just want to use the chat directly, you can go inside there or you can go over here. I actually prefer to use this because it's way nicer.

And then also cool thing about having an agent operating system, if you haven't already set up a mission control, is that you can see everything in one place. So with Gemma 4, we can preview over here and it looks really cool, right? And we can just have a look and see what we've created, what we've done. We ran some tests as well, as you can see right here, for like general day-to-day stuff as well.

And it was not bad. It's not amazing at writing. That's one thing that I'll be honest about as well. So the SOP, how to pair them up.

You point Hermes at Gemma 4. You can open Hermes and ask for something. You can connect it to your Obsidian vault locally so that it has a memory. You can put its tasks on a schedule.

So if you have tasks that you do daily, for example, like research or content creation, you can schedule that inside Hermes. And then you could ask it to build something that could be a mini app or a tool or that sort of thing. You can see it also uses multiple languages as well. So it can use multiple languages.

It can do marketing tasks like you can see here and it explains things pretty simply as well. It's just a laptop size model, which is pretty amazing as well. Now, if you're wondering, okay, how does it do on the benchmarks? So you can see here that it uses advanced reasoning.

It's laptop ready. So it can run with just 16 gigabytes of VRAM. It's open source as well, which is great. And you can use AI agents locally now as well.

And you can see how this performs versus the other two models, the 27B as well. Not only that, but you can also run Gemma 4, not just with Hermes, but with OpenCore, with Codex, with Cloud Code, with OpenCode as well. So you can run it with all your AI agents and everything you do agentically as well, which is pretty cool. So just to recap, Hermes is the body that's running 24 seven.

Gemma 4 is the brain. Together, they can automate tasks. They can schedule stuff. You can manage all the models.

You could use this as a sub-agent and you can also use it offline as well. So that's really cool. Something that's changed now is, let's say, for example, you're going to take a flight somewhere. You can use Hermes offline now with Gemma 4, runs on a schedule.

And if you don't have Wi-Fi, it's no problem. So that's basically how to use it, how to set up, how to get it working, et cetera. If you want to get more training on this sort of stuff, feel free to get the AI Profitable Boarding. There's an amazing community where you can ask questions, get help, get support in real time.

Very active community where we answer your questions daily. There's always people online as well, 24 seven, which is great. Inside the classroom, you can get access to all my new training. So if you're a complete beginner, you can go from beginner to expert in just six weeks with AI.

You also get my new daily tutorials and video updates like you can see. So we update this daily with new roadmaps and systems and prompts and 30 day plans on implementing this stuff. You can get the agent operating system with the video tutorial and a zip file that you can just install for all of this, plus the prompts on how to set it up. And additionally, you get the whole memory system, which we use with Obsidian, which is really cool too.

You also get four weekly coaching calls where you can jump on calls, ask questions, get help and support in real time. And inside the map, you can also meet people in your local city who are using AI agents just like you, which is great as well. And this is all inside the AI Profit Boarding. Link in the comments description or just go to the AI Profit Boarding.

Gatsby says, thanks for watching.

More episodes

Browse all episodes →