AI News Today
← All episodes
Episode 20 · June 4, 2026 · 11:25

New Google Gemma 4 Update is Insane (FREE!) 🤯

Gemma 4 12B Just Dropped: Free Agentic Open-Source Model You Can Run Locally

Google has released Gemma 4 12B, a new free, open-source, agentic model you can access now and run locally, and the creator demonstrates building real apps inside an agent operating system by plugging Gemma 4 into Hermes Agent. Examples shown include a mouse-follow animation app, a color palette designer, a Pomodoro timer, generative art, a simple game, a website, and a wallpaper, emphasizing improved usefulness versus earlier Gemma versions. The script explains how to get Gemma 4 via Ollama (recently updated), or use free Gemma 4 API options like a 26B model via OpenRouter if you lack powerful hardware, and highlights core specs such as 16GB memory and a 256K context window along with benchmark claims around 77% reasoning and ~72% coding.

00:00 Gemma 4 Drops
00:32 Demos Built in Hermes
01:35 Get It via Ollama
02:20 What Gemma 4 Is
02:56 Using It in AOS
03:50 Goldie Pocket Genius
06:50 Setup and Integrations
07:33 Use Cases and Limits
09:08 Benchmarks and Multimodal
10:04 Recap and Community
11:23 Final Goodbye

Full transcript

Gemma 412B from Google just dropped. This is their new agentic open source free model that you can get access to right now. I've been building with it inside the agent operating system. It's creating some pretty cool stuff, which I'll show you in a second.

I'll show you exactly how it works. You can also plug it into your AI agents, which I'll come on to later. And this is a free model that's designed to be agentic. Now, if you're wondering how it performs, what are the benchmarks like, we'll come on to that in a second, but you can see it does pretty well.

This is some of the stuff that we've built with it. As you can see right here, if you want to have a look, as I've been using it inside the agent operating system, basically plugged in Gemma 4 directly into Hermes agent. And you can see some of the cool stuff that it's built right here. For example, we created this amazing app where if I move my mouse, the animations follow it.

You can see here, we created this cool color palette designer. We actually built out a timer as well, which seems to work pretty nicely. One of the things that I've always found with Gemma 4 previously is, you try and build something with it, but it's not that useful, or it's just not as smart as it previously was. But this is actually building useful, real stuff that you can do this locally for free.

And this is with the brand new model. So it is a lot more powerful. It's a lot more agentic than anything else I've ever seen from the Gemma 4 models previously. Here's another example.

So this is a game that you built. I've shown you examples, apps, tools, games. You can even build out websites with it. We've got, for example, this cool little wallpaper that we created with it.

It's very visual as well. That's what I like about it as well. So everything we've created, it looks cool. It feels good.

Let me show you this as an example right here. And if you want to get access to this, you can go to Olama, and Gemma 4 is waiting for you right here. And it's literally just been updated four hours ago. So this is a brand new update.

This is the new version from Gemma. Gemma 4 12B is the new model that just dropped. So this is Gemma 4B, by the way. The thing I would say here as well, is if you don't have a setup that's as powerful as this, you can get Gemma 4 APIs for free and then plug those into your AI agents.

So let me show you an example of that. If you go over here and we type in Gemma 4, you can see that we have some options. So you can get from open router, and then you can just plug that into your Hermes agent. So even if you don't have an amazing local setup, it's okay.

I'm going to walk you through exactly what it means, how it works, et cetera, in a second too. And you can see the benchmarks right here. So let's talk about Gemma 4 12B, what it is, what it means, et cetera. So this is a lightweight model that's designed to be agentic, right?

So it can run on a normal laptop, yet it reasons near the levels of models twice its size. Now everything below it built or wrote itself live through Hermes agent today. So the stuff that we've built down here, we created with Hermes agent and Gemma 4B. So it's a 12B model, 16 gigabytes memory to run it, K context window, and it's open source and free.

That means it's free, it's local. You can run it locally, you can run it, et cetera. And these are some examples of apps that we've built. Now, the way that I did it is basically you can just go into the agent operating system, change your model to Gemma 4, and then from there you can start chatting with it.

And then the cool thing about this is if you're using it inside an agent operating system, everything that you've created, you can view and preview inside the workspace as well. So everything is saved. Like for example, when we were creating stuff with minimax, we can see everything that we've created here and we can just preview and go back to projects we've previously created, which is super useful. Here's another example.

So this is a game that we built out. Everything seems to run pretty smoothly as well. I also like the way that it can do this sort of art. Look at this.

How cool is this? This is pretty crazy. Well, you can build with this. And then it also did some like reasoning tests with it as well.

Not bad. It can write decently. Bear in mind, like this is not gonna be at the same level as something like Claude Opus, of course. If you're expecting like Claude Opus 4.8 levels, you're not gonna get it with this, but it is really cool.

And you can see some of the examples we've built with it. So let's talk about this. I've built out a framework called the Goldie Pocket Genius. This is exactly how you can use Gemma 4 locally, running for free and how to get the most out of it.

So basically what is Gemma 4 and 12B? We have released many different models in the past. It's a whole family of different Gemma 4 models. You can get different sizes, et cetera.

12B is a super lightweight version that's designed to run locally. And it's more gigantic as well. So it's designed for gigantic. It's 256K context window.

And also it can plug into your AI agents like Hermes. And the cool thing is, well, if you're running a local model, obviously your data never leaves your computer and you can run it unlimited. So you could be like working offline, for example, you got no wifi, no problem, you can still use an AI. And so the other cool thing is you can drop it into Hermes.

Let me show you how to do that. So basically you go into your terminal here and then you would run the command Hermes model. And then from here, you can select Olama if you want to run it locally. So you can run Olama with it.

Or if you want to run a free API with Gemma 4, you can actually run 26B, which is free with open router. That is not the new model, but if you want a free API, that's another option that you can use. And then you still get the power of Gemma 4. So it's free to run, 16 gigabytes, runs on a laptop, 100% private and offline.

Before, obviously, if you weren't running local models, what would live in someone else's data center and you just kind of have to pay for the API or for the tokens. Now you can use something like Gemma 4 locally. And basically the way that I look at this is that Google just put like a little AI on your laptop. So every AI agent like Hermes is a worker, right?

To think it needs a brain, the model that reads, reasons, and writes. Until now, that brain almost always lived in the cloud. So every thought would require tokens, resources, and your private data left your machine to actually use it. Gemma 4 fits that because Google built it to be small enough for your laptop, but smart enough to rival the giants and gave it away for free for anyone to use and keep.

The Gemma family has already been downloaded over 150 million times. So this is a version that's a bit more lightweight and easier to use. And also it's designed for agentic reasoning. So the agent is a worker, the model is its brain, and Gemma 4 is a free brain that you can plug in.

Now, why does it fit on a laptop? Well, older models were basically heavy because they bolted on clunky extra parts to handle images and sound. Gemma 4 is built differently, understands words, pictures, audio natively in one unified brain without the bulky add-ons. So the bulky add-ons have actually been stripped out of it.

And that's a trick that shrinks it enough to run on a normal machine. It was also tuned for agentic reasoning. And that's a fancy word for saying it's good at planning and it's good at doing multi-step jobs, not just chatting. And that's exactly what an AI worker, like for example, Hermes or OpenCore or anything else that you use, needs.

And it holds a small book in its mind at once, basically as a 256K token context window so it doesn't lose the thread on long agentic tasks. Now, I've actually created a framework called the Goldie Pocket Genius, which is you can download it, then you can run it, and then you can actually unplug it. So you can turn off the Wi-Fi and it keeps working offline private. And this is a quick five-minute setup, really.

The first thing to do is download it. If you want to download it, make sure that you have Olama and the latest version of Olama downloaded. Then from there, you're gonna make sure you have Olama running. And then from there, you can copy any of these commands to run it inside each one.

So you can run Gemma4 inside Hermes agent, OpenCore, Codex app. Now, would I run it personally inside all of those? Now, I would say the Hermes agent is pretty good because it's lightweight. It doesn't use up many tokens.

It's designed to be token efficient. If you use it inside something like Cloud Code, probably wouldn't be that powerful to run inside there, but you would get more out of it inside Hermes agent or just running it directly. You know, with Olama as well, it's a good option too. And then you can see the options here.

So 12B is the new version that's just dropped. So let me give you a use case and an example of this. You could, for example, if you're working on client work and you're running projects for a client with sensitive data, well, you can use a local model that's private. And it's free as well.

So you could even work, for example, on a plane using AI and it would still run. Now, let me be 100% honest with you. You know, like this is a local smaller model. So for like super hard reasoning tasks, or for example, like coding tasks, would you use Gemma for?

Probably not, unless you have an amazing setup. But if you wanted something like, for example, like just a day-to-day task or being on a landing page or a quick mini app, like I've shown you before, I mean, these are all the things that we've built with it. And you can see like we can generate like pretty amazing animations. We can create games with it.

We can do design jobs. We can create mini apps like the Pomodoro app. Like, so for example, if we open this up and we test out, it works perfectly. And this was fully built with Gemma for, which is pretty cool.

And then here's another example of a game that we built out. And it is very visual. It looks really cool, et cetera. So the way that I've been using it is inside Hermes, as you can see here.

So we said, what model are you using? It's using Gemma for. And then the cool thing about this is if we go over to the workspace, everything is grouped into all the stuff that we used and tested it for in one place, which just makes it easy to come back and see what we've built previously. And also like, I like this because I can test between, okay, what have I built with different projects and different models?

And then which one compares best for different tasks. So for example, like for Minimax, Minimax is actually pretty cool for like voice agents and for example, generating videos. Whereas for example, Gemma for, it's great for running locally and running it for free. Different use cases for different models.

Now you might say, how smart is it really? So here's where it actually scores. On a tough general reasoning test, it hits 77%. On a real coding benchmark, it hit about 72%.

Those are numbers you'd expect from models twice the size and it does it using half the memory on a simple laptop. On top of that, it can also read images. It can listen to audio, not just text. So for example, you could hand your AI agent a screenshot, a photo, a voice note, and it could actually handle it, which is pretty cool.

Now, some people are gonna say, good AI only runs in the cloud, you need a super computer, but this is a free lightweight model. Other people say, well, local AI isn't that good, but actually you've seen the examples of what we've created right here, it's pretty cool. Other people, I see it all the time, they're like, you know, I don't wanna pay for a subscription or no tools, et cetera. So this is what Gemma for solves because you can run it locally, it's free, and you can use it unlimited.

And then other people say, well, I'm too late, everyone's ahead of me. Honestly, this literally just dropped today. So just by watching this video, you're ahead of most people. So just to recap, every AI agent needs a brain and that brain used to live in the cloud, but now you can use something like free, locally, plug it into Hermes agent.

I've shown you how to do that. You can just run Hermes agent inside your terminal. Once you've downloaded the model, it builds some pretty awesome stuff and we've plugged it already into our agent operating system. If you wanna get the agent operating system, by the way, we have OpenCloud, Cloud, Hermes, Gemini, Antigravity, Codex, FreeCloud Code plugged in, along with some really cool workflows for SEO, video agents, notebook glam, et cetera, even a full Kanban board for building a team of AI agents with this stuff.

You can get it inside the AI Profitable Boardroom. This is my AI automation community that shows you how to learn, grow, and save time with AI automation. There's 3,300 members inside there, which means there's always people online you can ask for help inside the community and ask questions, et cetera. It's a very active community with lots of people asking cool questions.

You get access to all of my best trainings, including the new daily updates. So we had new daily guides here with new tutorials and step-by-step guides, as you can see. And then we also have the agent operating system with the zip file that you can download and install right there, and a video tutorial on how to use it, plus the last update date. You get four-week coaching calls where you can ask questions, share your screen, meet other cool people doing similar stuff.

And also inside the map here, you can meet people in your local city who are using AI agents just like you, which is pretty cool as well. So feel free to get that link in the comments description or just go to the AIprofitboardroom.com. Thanks for watching. Cheers, bye-bye.

More episodes

Browse all episodes →