AI News Today
← All episodes
Episode 85 · July 16, 2026 · 08:42

New Google Gemma 4 Update is Insane (FREE!) 🤯

Full transcript

Google Jemma 4 has a brand new free update. This is a local model and it is more powerful than ever. Here's an example of a website we actually built with this free local model, absolutely wild. And it's really easy to use.

Now you might be wondering, okay, what changed? I'm gonna walk you through exactly what's changed today. It's massively improved in terms of the benchmarks and how it's performing and everything else. You can see some examples right here.

If you wanna get access to it, you can get it on Hugging Face. And I'm gonna show you some examples of what we built with it first. Again, this is free, it's local, easy to use, can create some beautiful website designs as you can see. Here's an example of a quick app that we built out, kind of like a mini app.

So you don't just build websites with it, you can do whatever you want. And here's an example of a game that we actually built out with it as well. So Jemma 4 has become a lot better and faster since this new update. I'm gonna walk you through exactly how to use it using something I call the Limitless Local Machine.

So Google just shipped a big Jemma 4 update with Wired Deleted version into our agent operating system as a free local brain. And it's really, really powerful stuff. If you're wondering how to use the AgentOS, link in the comments description or go to theairprofilebomb.com to get it. And if we go over to the local section here, you can see that our brain is plugged in as Jemma 4 into the system and we can build a code wherever we want.

Plus we can save the creations that we've made over here too. Now you might wonder, okay, what has changed? What's different now? So basically Google has just made their existing Jemma 4 model much better than before.

And what this one does is use MTP drafters, which lets the model draft its own next words at a time. What does that mean in reality? It just means that it moves way faster and that speed up you can actually use inside Odalama, which is a free tool you can run in your Mac, right? Now, why does that matter for you?

Because Jemma 4 is free, it's private, it's offline as well. You can run locally and it's fast enough to really build some amazing stuff, as you can see right here. So these are the net improvements overall in terms of how it's performing. So before it would work like one word at a time and it would like have a word, wait, word, wait, et cetera.

Right, that makes it much slower. With this system, it moves way faster because it uses MTP drafters. So it guesses ahead as an AI and then it checks the works up to three times faster. And we ran it on a Mac using this system and the builds that we created and it looks super nice.

So why would you use this? For example, a lot of people worried about tokens right now. And also if you're working offline, then this is a way to use the power of an agent operating system or any agent you use. So even for example, Hermes, you could run offline with Jemma 4 and now it's much faster to use as well.

It's also improved in terms of benchmarks. So Jemma 4's 31B jumped from 20.8% to 89% on Amy, which is pretty amazing in itself in terms of benchmarks. And this whole brain lives on your machine. So you want to think about it like this, like cloud AI is kind of like a phone call.

So your question travels to someone else's computer, it gets answered there, then the answer travels back. And there's kind of like a meter because you're using their machine and you only have a limited number of tokens. A local model is different because you download the entire brain once. For example, like 12B is only seven gigabytes and then it can run on your Mac or your computer.

And there's no call, there's no meter, nothing needs a machine. You can ask it once or you can use it hundreds, if not thousands of times, and it doesn't cost you anything because it's a free local model. And then you can unplug the internet or you can switch off your wifi and it still keeps working as well because it's not going to a cloud. And it's just a really smart way to run local models.

Also Jemma 4 is one of the few local models I've seen that actually works agentically. And to get started, you just need to install Alarmer, pull the latest Jemma 4, test it out. And then we've actually put it inside our agent operating system so that it's ready to go whenever we need it. So if you look at the old way of using something, which is, for example, paying for an API, you know, every question, you got limited tokens, you get rate limited sometimes, you're prompted to go into the cloud.

If the wifi goes down, your AI is gone. And then also you have no control over the model. With a new system, we're using the limitless local machine. There's no meter, so you can use it all day, every day.

There's no rate limit. Nothing leaves your machine. It can work on a plane. You can have the wifi off and also you and the model as well.

You might also say, well, is my setup powerful enough to run this? So Jemma 4 actually comes in lots of different sizes, but the 12B model, which is eight gigabyte, works totally fine with our Mac Studio. And we've tested out, as you've seen with three different examples. This way as well, like instead of you worrying about the models, you have the machine, right?

So if you set up Jemma 4 inside an agent OS, you know, we all saw this with Fable 5, we saw GPT 5.6 in preview, so we couldn't use it for two weeks. The main point here is that if you have the system and you don't worry about the models, that's what's the most important thing, because the models are going to change. Like local models, they're going to improve. They're going to get better and you can swap them in and out, but using this system instead with the agent OS, it doesn't matter what comes out, what changes tomorrow, we still got a system that improves every single day and gets better every single day as well.

You also might say, why would you use a local model inside an agent OS? Well, number one, it reduces the amount of tokens you use. Number two is free. And number three, it can run side-by-side versus your other agents.

So you could have, for example, like Codex running, which we've set up inside the agent OS, and then that could actually orchestrate your local models, which works really nicely, and it's plugged into the memory system, so you're good to go from there. Now, some people say three local models are way behind. I would say Jemma 4 is one of the few ones that actually creates good stuff, as you've seen before. You also might say, okay, setting up a local model, that sounds like a lot of work, but it's just one terminal command.

So you just download Olama, and then you run this command, and you can run Jemma 4 with it. And then also you might say, okay, a local model adds nothing, because I've already got a subscription. But when you have them working together, you can have Codex, or you can have your frontier models, like Fable 5, orchestrate your local models. So you reduce the amount of tokens, so you just have the frontier model as a brain, and then your sub-agents can run with local models instead.

You also might say, okay, this stuff sounds technical, like difficult to use, but actually we've got over 200 pages of people using agent operating systems and local models, and all these members are members of our community, the Aeroprofit Boarding, and they're getting awesome results. I'm non-technical, I've built out the agent OS. You don't need to be technical to use this sort of stuff. You just need to get started.

So you can install it, wire it into the agent OS. Then the hard sort of like time-consuming tasks that don't require a lot of intelligence, you can delegate to Jemma 4, and then it's free to go. And if you want to get this set up with the agent operating system, you can get that inside the Aeroprofit Boarding. We've got the local machine inside there.

We have all sorts of builders. If you like free stuff as well, we've got OmniRoot and HY3Coder built in there. These are two different coding agents that you can use for free as well. And then also you get the full sit bar to install.

You get weekly coaching calls where we set up your machine together. There's 4,000 members in there, so you get daily tutorials and help and support 24 seven, because there's always someone online, and we've had over 200 pages of wins in there. So what have you learned today? You've learned why you shouldn't just rely on APIs, because it can get expensive and use up a lot of tokens.

Also, what the latest Jemma 4 update means, basically that it can run three times faster, how to install it, so you just run this one-term command once you allow my setup, and then how to use it inside an agent operating system, and why to use it. So it's offline, it's private, it's free, and you can get your agents, your frontier agents to actually orchestrate them for you. So if you want to get all of this set up, get it inside the iProfitboarding, link in the comments description or go to theiprofitboarding.com. Inside the community, you can ask questions, and I answer them personally with a new daily video tutorial.

Plus there's always people online to help you whenever you need to. Inside the classroom, you can actually ask questions, and you can get all of our new training. So we actually have new daily updates based on the new and advanced stuff. We actually have lots of training on Jemma 4 and local models as well.

So if you type in Jemma 4 inside the search bar, you can see we have loads of different trainings on Jemma 4, how to use it, how to get the most out of it, et cetera. And then also, if you want our agent operating system, that's over here, we update it regularly. You can get the videos you saw on how to use it, and you can get a zip file. If you actually just want to build your own from scratch, we have a full one hour course on how to do that here.

And then inside the calendar, you can drop a weekly coaching goal, share your screen, meet other members, et cetera. And you can meet people locally in your area, running local models and the agent operating system just from the map. So see inside there. And of course, you can DM me directly, link in the comments.

More episodes

Browse all episodes →