AI News Today
← All episodes
Episode 1 · April 4, 2026 · 10:05

Google's Gemma 4 AI Just Changed Open Source Forever

Google Gemma 4: The Ultimate Free & Local AI ModelGoogle's new Gemma 4 changes open-source AI forever by offering frontier-level performance in a completely free, locally runnable model. Discover how its massive context window and native agentic features enable private, cost-effective automation for your business.00:00 - Intro: Gemma 4 Changes Everything01:05 - New Apache 2.0 Licensing02:08 - Model Sizes & Local Performance03:16 - Multimodal & OCR Capabilities03:58 - 256k Context Window Explained04:21 - Native Agents & Tool Use05:50 - Efficiency vs. Large Models08:50 - Final Summary & How to Access

Full transcript

Google's Jemma 4 just changed open source forever. So Google just dropped Jemma 4 and this one is different. This is a new local model you can run for free with Google. The 31B model, their biggest, just hit number three on the Arena AI open model leaderboard worldwide.

Out of every open model on the planet and it's completely free to use commercially. On maths benchmarks, AIME 2026, it scores 89.2%. On science reasoning, 84.3%. On competitive coding, 80%.

Those numbers were what frontier paid models were hitting a year ago. But here's the part that actually matters for your business. You don't need to pay OpenAI or Anthropic to run this. You run it yourself on your own hardware.

Your data stays with you and it's built from the same research as Gemini 3, Google's top proprietary model. If you want to get access to this, you can go to Alarma and install it. We've got tons of training on that inside there apart from boarding. And they took the same tech powering their premium product and gave it away for free under a license that lets you do whatever you want with it commercially.

That's a big deal. Let me explain why. For the last two years, businesses wanting to use Google's AI models hit a wall. The license had restrictions.

Legal teams flagged it. Enterprise companies said, no, no, no, no, no. Compliance teams flagged edge use cases and capable as Gemma 3 was, open with asterisks isn't the same as fully open, my friends. Gemma 4 eliminates that fraction entirely.

So you've got a shipped version of Gemma 4 under a standard Apache 2.0 license, the same permissive terms used by the rest of the open way or ecosystem. No custom clauses, no restrictions on redistribution or commercial deployment. Hugging face co-founder Clement Delang called the licensing change a huge milestone. And he's right because this changes what small businesses and solo operators can actually do.

You can now build a product on top of Gemma 4. You can use it for automating your work. You can host it privately. You can run it on your own machine with your own data.

No monthly API bills, no per token costs, no sending sensitive data to a third party. Now, four models dropped and you need to know which one is for you. So there's a tiny E2B and E4B model. These run on phones and small devices.

Then there's a 26B mixture of experts and the 31B dense model, the big one ranked number three in the world. Early testing shows the 31B running at about 10 tokens per second on consumer hardware. The smaller 26B hits over 40 tokens per second. And for most business tasks, writing, research, summarizing, drafting, that's more than enough and more than fast enough too.

If you're running a content agency and you don't want to send, for example, client briefs to opening our servers or any other server, you spin up the 26B model on your own machine and process everything locally, privately, fast, and free. And bear in mind, you can use this inside LM Studio. You could run it inside Hugging Face. You could run it inside Olama.

You've got many different options for actually hosting it and using it. The 31B fits on a single 80 gigabyte NVIDIA H100 and Qantas sized versions run on consumer hardware. So if you've got a decent GPU setup, you can run the top tier model yourself. And it's not just text, right?

Every single model in this family is multimodal out of the box so they can handle images, text, video, and the edge models and native audio, automatic speech recognition, all on device. The models handle object recognition, reading PDF documents, and OCR. So if you're running, for example, an e-commerce business and you want to process supplier invoices, contracts, or product images, Gemma 4 can handle all of it offline without sending anything out. And the Vision Encoder also supports image token budgets from 70 to 1,120 tokens per image, lower budgets for quick classification, higher budgets for detailed document parsing.

The context window matters here too. So you've got the edge models which handle 128 tokens, the larger models which can handle up to 256 tokens. That means you can drop an entire document, a full client report, or a long email thread into a single prompt and get a complete answer. A 256K token context window means you can paste in mumps of customer emails, your full product catalog, or a 300 page contract, and ask your questions just like a human assistant who can read the whole thing.

And the agentic side is where it gets serious for automation. So function calling is native across all four models. Unlike previous approaches that relied on instruction following, Gemma 4's function calling was trained into the model from the ground up, optimized for multi-turn agentic flows with multiple tools. What that means essentially is you can plug it into OpenCLR and it'll actually work.

That's not a small thing. What it means is that Gemma 4 can actually use tools reliably, it can call an API, it can check a database, it can run a search, it can take a sequence of actions without breaking. If you're building any kind of AI automation, lead gen workflows, content pipelines, client reporting, Gemma 4 can power the whole thing privately on your own hardware. And here's the thing about AI automation right now, the businesses winning are the ones who figure this stuff out, who figure out how to run these APIs and run them locally without paying for everything, because that reduces costs, which increases margins, right?

Now, if you want a 30-day roadmap for building private AI workflows using models like Gemma 4, automating your lead gen content or client delivery, and you want coaching calls with people who are already doing it, come join the AI Profit Boarding. 2,700 business owners inside there right now. We've got step-by-step tutorials on setting up local AI, running open models, and building the kind of automations that get you more clients and cut your costs dramatically. Daily tutorials, four coaching calls a week, and a prompt library built around these exact tools, plus a member map so you can connect with other business owners running open-source AI setups near you.

Link in the comments description or go to theairprofitboarding.com. Now, the performance numbers. Google claims both larger models out-compete models up to 20 times the size on the Arena AI benchmark, 20 times their size. The 31B is beating models at 600 billion parameters on certain benchmarks.

There's the efficiency story as well here, because Google squeezed Gemini 3-level intelligence into a model small enough to run on your own setup. The 31B currently sits at rank three on the Arena Open AI model leaderboard. The 26B is at rank six. And the honest caveat here is, compared to the biggest Chinese open-source models like DeepSeek, Gemma 4 doesn't keep up with that weight class, right?

It's really designed for local devices and things like phones and that sort of thing. But DeepSeek's largest models are still ahead here. But having said that, this is more efficient and smaller to use. And for most business use, like writing, research, automation, document processing, you don't need to beat the biggest models in the world.

You need a model that's capable, fast, private, and free. And Gemma is all four of those things. The on-device angle is actually massive for the next 12 months, because the Edge models are up to four times faster than previous Gemma versions that use up to 60% less battery. Early tests on the E2B model show a 5.5x speed up in processing user input, and up to 1.6x faster response generation on ARM CPUs.

Android developers can prototype agentic flows in the AI core developer preview today. So business building apps are about to have access to a genuinely capable AI model that runs entirely on the user's phone, which is pretty insane. So that means no server costs, no latency, no data leaving the device. And think about what that means, for example, if you're like running a client-facing app.

You know, a consultant who runs their entire client intake process through an AI app on their phone, a freelancer who uses a private AI assistant that never touches a cloud server, that's where all of this is going. Now Gemma has accumulated over 400 million downloads, more than 100,000 community-created variants since its first release, and that's not really a vanity number. That means there are 100,000 fine-tuned versions of previous Gemma models already in the world, right? And built by researchers, built by businesses and developers who customize the base model for their specific use case.

Now with Gemma 4's Apache 2.0 license, that number will probably explode and every industry is gonna have its own version. A Gemma 4 fine-tuned version on legal documents, one may be trained on customer support transcripts, one built specifically for e-commerce product descriptions, all of them free to use, free to deploy, and free to modify. You can train and adapt Gemma 4.2 using Google Cloud, Vertex AI, or even your gaming CPU, a standard gaming PPC, not a data center. That's what changes when you use these sort of models.

And also, I was running it on my Mac Studio earlier today. It was pretty chill. It worked nicely with OpenCore. It was pretty easy to use and it was free, completely free to set up and quick to set up too.

So let me pull the full picture together here. Google just gave away their most capable open model family ever, free commercial license, runs on your own hardware, beats models 20 times its size, handles text, images, video, audio, 256K context window, and native AI agent tool use. And it's available right now on Hugging Face, Kaggle, and Olama. Google DeepMind CEO, Demis Hassabis, called this the best open models in the world for their respective sizes.

The businesses that move fast on this are gonna have a real advantage, and private AI that costs nothing to run is pretty amazing. Agents that work without breaking, document processing that stays in-house, lead gen and content workflows that run completely on your machine for free. The gap between what big companies can do with AI and what small businesses can do just got smaller, and that's the story here. If you wanna know exactly how to set up Gemma 4 for your business, whether that's local inference, building private automation workflows, or using it to process documents and generate content without API costs, the AI Profit Boardroom is where we walk through all of it, step by step.

We've got members already running open source model setups for client work, content pipelines, lead generation. You get four weekly coaching calls and everything you need to win with this stuff. Link in the comments description, or just go to the AIProfitBoardroom.com if you wanna get access. Thanks for watching.

More episodes

Browse all episodes →