AI News Today
← All episodes
Episode 20 · June 4, 2026 · 11:55

Google's NEW Gemma 4 Just Changed AI Forever...

Gemma 4 12B: Google's New Private & Free Local AI Model

Discover Google's latest Gemma 4 12B model, a lightweight yet powerful multimodal AI you can run privately on your own laptop. We break down its Mixture of Experts architecture, impressive benchmarks, and how to set it up for your AI agents for free.

00:00 - Intro to Google Gemma 4 12B
01:03 - Multimodal Intelligence Explained
02:24 - Hardware Requirements & Size
03:44 - Mixture of Experts Architecture
04:39 - 256K Context & Privacy Benefits
05:42 - Performance Benchmarks vs Larger Models
07:04 - How to Run Gemma 4 Locally
09:12 - Free API & Agent Integration
10:31 - Recap: Is Gemma 4 the New Standard?

Full transcript

Today, I'm going to run you through a Gemma 4 from Google. This new model that just dropped today, it is Gemma 12B. And this is a local free model you can run and it can plug into all your AI agents. You can see some cool stuff that we actually built with it right here.

And what's interesting about this is local. This is free. This is private and you can plug it into whatever you want. So for example, you can get it for free and then you can plug it into something like Olama, or you can plug it into Claude or Codex or that sort of thing.

It's actually designed to be a super lightweight model. So it's not going to be like the most powerful thing you've ever used, just to be 100% honest with you, but you can build stuff with it. As you've seen today, I've already built out games, tools, very visual stuff as well, which is pretty cool. It seems to be super fast.

And also you can actually get a free API for this. So even if you are not interested in running local models, maybe you don't have the right setup, et cetera. I'm going to show you exactly how you can use Gemma 4 for free as an API. And then you can also plug that into your AI agents or whatever you want to use.

Now, this is basically a local model designed to be multi. So this is your new announcement, which you just dropped today. This is designed to be a high performance, multimodal intelligence model directly to your laptop. And it's also designed to be mobile first.

So it can be pretty efficient for running on mobile and it's got advanced reasoning. And one of the interesting things about this is performance of benchmarks. So what that means essentially is that despite it being quite a small model, for example, the context window is only 256K and it's a 12-bit model. Despite doing that, it's actually operating at similar benchmarks to models twice its size, which is quite remarkable in itself.

So there's five key things that make this model unique. Number one is its novel unified architecture. So it doesn't have multimodal encoders, right? It's actually just in one single LM.

Also it's advanced reasoning performance. So it's pretty good at advanced reasoning. You can use it on a small laptop. So even like a 16 gigabyte VRAM laptop, you can use this on.

It's open source and it's released under an Apache 2.0 license. And additionally, it comes equipped with multi-token prediction to reduce. Now you can see the benchmarks right here in terms of how it performs versus other models. So this 27B and this Gemma 4B, bear in mind like Gemma 4 itself has been downloaded like 150 million times.

So it is a very proven model that a lot of people are using right now, which is pretty cool. And so let me talk you through why it is unique, what it means, how it works, etc. So Gemma 4, 12B fully explained. Number one, the size of its brain.

So 12B, the B means billion. A brain size is counted in terms of how many tiny knobs, I guess you could call it, it has to learn with. So Gemma 4 has 12 billion of them. That's small enough to fit on a 16 gigabyte laptop.

There's an even lighter eight gigabyte version too, but big enough to genuinely be useful. So you want to think of it as like a compact, but useful rather than oversized AI brain and a brain that you can plug into your AI agents. So let me give an example. We've already been testing it with Hermes agent and it created some pretty cool stuff.

Like it created some nice websites. As you can see, you could, for example, create these cool tools, these cool visualizations. You can see how it's generated this report here. It creates some nice keyword research as well.

And you can see it's beautifully designed. You can get quite a lot out of it. If you have the good, if you have good skills in place, you can see how it can do this. And then for actual just general project, projects in general, you can see some examples of stuff we've created right here.

It can create these visual apps, which are really cool. Create a color palette, a Pomodoro timer, a snake game. This was another sort of brick breaker game as well. Additionally, this is a reflex game that generated that measures your reflexes.

Wallpaper, which was pretty interesting too. And then we tested it on some general content and responses and that sort of thing. One interesting point here, it is multilingual, so it can speak multiple languages too. Also, one thing to note here is it's a mixture of experts, which is why it's fast and smart.

So the clever trick with Gemma4 is that in Gemma4's own words, instead of one brain trying to know everything, the AI agent is filled with hundreds of tiny specialized mini experts. One's a maths whiz, one's a history buff. So when you ask a maths question, it's only going to use the maths part of that model. And that's why a small model can perform so well and punch above its weight because it's not wasting energy, waking the wrong parts.

And pretty well explained. That was an explanation directly from Gemma4 on how a mixture of experts models work. Now, one unified brain, what does that mean? It's got one brain for words, pictures, and sound.

So older models bolted on clunky extra parts to handle images and audio. Gemma4 actually understands all three natively in one brain with the bulky add-on stripped out. So that's a design that shrinks it enough to run on a normal machine. And it means your agent can now handle a screenshot, a photo, a voice note, not just text.

Also, it's got a 256K context window. So what that means is it can basically hold roughly a small book in mind at once. Bear in mind, there's a lot of APIs that people still use that are still 56K and they pay for them. But this one you can use for free because you can run it locally.

The other cool thing here as well is that because it's running locally, if you jump on a flight or if you don't have wifi, you can still use not just Gemma4, but also your AI agents. So if we plug this into OpenCLR or into Hermes or even into Cloud Code, if you wanted to, which I don't recommend, but you could, then you could actually use this offline. You could have it running 24-7 and it doesn't matter if you've got wifi or not. And then also it supports 100 languages and it's Apache 2.0 open weights.

So it speaks over 140 languages out of the box and Apache 2.0 open weights simply means the brain is truly free and yours. So you can download it. You can run it offline forever. You can even use it commercially.

No account, no meter, no one who can switch it off, which is pretty cool. And this really like the moment the laptop caught up to the cloud. Now, is it the same performance as Cloud Opus 4.8 or something like that? No.

Absolutely not. But is it useful? Can you build cool stuff with it? Can you run it offline?

Is it actually usable? Absolutely. And so you might be wondering, okay, how smart is it really? So the benchmarks are just standardized exams for AI, same test, every model, et cetera, so you can compare how it performs.

You've got 77% for MMLU Probe, which is a brutal general knowledge and reasoning exam across law, medicine, maths, and more. It's a score you'd expect from a model twice its size. TENTS is scored 72% on live, which is basically testing real fresh coding problems, so it can't have memorized the answers. So these are new answers, not based on memory that it gets tested on.

And then also 26 bits of quality per size. The independent testers peg its answer quality near a 26 billion parameter model at half the size. That's the headline. It's really like a heavyweight answers model, but featherweight footprint.

And the context, you can read about a small book with the 256K context window. So what does this actually mean for you, aside from all the benchmarks that most people don't pay attention to anyway? So translation for everyday business work, drafting, summarizing, answering from your own documents, writing code, running agent jobs, it performs like a model you'd normally have to rent from the cloud, except this one is free and runs on your desk, so you've got heavyweight scores, featherweight footprint, and it's free. Next up, the way that I see anyone using this and the framework that I look at it is called the Goldie Pocket Genius.

So you can own it, you can run it and unplug it, right? So most people are going to use APIs, but if you wanted to set this up yourself, then you can own it, which means you can download Gemma 4 free onto your laptop. If you're wondering how to do that, you can get it directly from Olama already. So if you go to olama.com, make sure you've downloaded Olama.

And then once you've done that, you go to Gemma 4 and you can plug it in with one single copy and paste prompt inside your terminal using this, or you can plug it into your AI agent. So for example, like Hermes or OpenCore, like I was talking about before, we already have it plugged into our AI agent. And the other interesting thing about this is you could use it for other things. So it doesn't have to be like your main model.

You could have a really powerful model, like M4, the main model. And then for auxiliary tasks, you could have your sub-agents running Gemma 4 as well. And so you can download it, that's step number one, already shown you how to do that. You can point, for example, Hermes as Gemma 4 as its brain.

You can do that inside the agent operating system, I've already done it. And your AI worker now thinks on your own machine instantly. And so if you want to run it inside Hermes, you can just get this response right here. And then finally unplug it, right?

So if you turn off the wifi or if, for example, you're on a flight, something like that, you can still run it offline, privately and unlimited. There's no rate limits and no one can ever switch it off because it is running locally on your machine, which is pretty crazy when you think about it. And so you might be wondering, okay, what can it do? So here's an example of something that actually coded.

You can see it built out this game right here. It also created some live generative art, which is pretty amazing. When you look at this, like very interesting visual. It can reason step-by-step.

So we actually gave it a five blue logic puzzle and you can see that it answered it correctly first try, which is awesome. It can write like a novelist. Although again, I wouldn't say it's at the level of something like Quartzsonic or Opus, but it can write for you and it can speak many different languages, so Japanese, Arabic, Hindi, et cetera, it can strategize as well. So it can help you with marketing tasks.

These are all tests that we've run and also we ran it with Hermes agents. So how'd you do that? So you would point Hermes at Gemma 4. So for example, if we go to our model settings, we would just configure this to use Gemma 4.

You might be wondering, how'd you do that? There's two ways that you could run it locally. Again, if you don't have a good setup, no problem. You can still use it for free.

Let me show you how. So you can go over to Open Router and there's a few free models you've got with Gemma 4 and 26B if you're using cloud-based models with Open Router, which means you can use this model for free as well, which is pretty cool. Then you can ask it stuff in plain English. So for example, like schedule the content that I do every single day, and it can write real tasks to your workspace.

So if we have a look inside our workspace for Gemma, you can see all the cool stuff that it built out directly here, right? Pretty amazing stuff. And so it actually does stuff. It can actually build things.

It can actually write and create amazing stuff. It is interesting. Now, where would you use a bigger brain? I would say a laptop size brain is not a frontier giant.

So for the very hardest reasoning tasks, stuff working for hours, long horizon tasks, et cetera, you're not going to use Gemma 4, but that might be like 10% of your daily work. For the other 90%, you could potentially use Gemma 4 if you really wanted to. Now you might be saying, okay, good AI only runs in the cloud. You need a supercomputer, but this is a free model on a normal 16 gigabyte laptop that built everything that I've shown you.

You might say local AI is just a toy. It's not that good, but actually you've done pretty well on reasoning exams and real coding tests. And you might say I'll always be stuck paying whatever AI companies charge. But when you use a model like Gemma 4, you own the brain.

It's free. It's open. It's yours and no one can price hike it or cut you up. And also other people say I'm too late.

I'm too late to this stuff. Everyone's ahead of me. Actually, it's brand new. It just dropped today.

So that's basically it. To recap, Gemma 4 is 12 bit. It's a free, open, useful brain that you can download and run on your 16 gigabyte laptop. It uses a clever mixture of experts designed to be fast and smart at a small size, holds a small book in its memory for 56K, context window, it speaks 140 languages and scores like a model twice its size.

It reads images and audio too. It's not theory. We've tested it. We've used it with Hermes Agent and you've seen all of the things right here.

Now, if you want to learn more about how to use it, or if you want to get the agent operating system that we've set up with Gemma 4 and Hermes Agent, you can get that all inside the AI Profit Boardroom. Link in the comments description or go to the AIProfitBoardroom.com. Inside the community, you can ask questions. It's a very active community where you can ask questions, get help and support, lots of people helping each other.

There's always people online, which is great as well. And everyone's super useful and positive, which is awesome. If you want to get the agent OS system, plus a zip file to install it, we've got it right here, along with a video tutorial and a breakdown, and we just updated it today to reflect the new changes. We had new video guides inside here with tutorials and 30 day roadmaps and everything else you need to win with this stuff.

Inside the calendar, you can actually jump on four weekly coaching calls. So you can share your screen, ask questions on the call, meet other cool people using Jemma 4 and other AI automations. And inside the map, you can actually connect with people in your local area who are using AI agents like you to meet them, hang out, et cetera, meet other cool people near you who are using AI agents just like you. So feel free to get that link in the comments description or just go to the AIProfitBoardroom.com.

Thanks for watching. Cheers. Bye bye.

More episodes

Browse all episodes →