AI News Today
← All episodes
Episode 1 · April 20, 2026 · 06:39

Claude Code Local DESTROYS Claude Code?

Run Claude Code Locally for FREE: Offline & Private AI Agent


Learn how to set up and use Claude Code Local to run powerful AI agents entirely on your machine for free. This guide covers offline usage, model switching between Llama and Gemma, and even controlling your agent via voice or iMessage.


00:00 - Intro to Claude Code Local

00:34 - Running Claude Code Offline

01:47 - Comparing Local Models (Llama, Gemma, Qwen)

02:33 - Local Performance & Private Data

03:23 - Fixing Tool Call Reliability

04:19 - iMessage & Hands-Free Voice Control

05:14 - Hardware Requirements (M1/M2/M3)

Full transcript

So today we're going to be testing out a new open source project called Claude Code Local that allows you to run Claude Code locally for free. You can also get it to speak to you as well, which is really cool. And you can run this with local models, like for example, Gemma. Gemma 4 is a free local model, same with Quen 3.5 and Llama as well.

And you can switch between them, right? So for example, if you want like something that's super powerful, then you would use Quen 3.5. This will be really powerful. You could also switch to Llama 3.3 if you wanted to.

And if you wanted a local way to use Claude Code really quickly, then you could use Gemma 4. They have a cool option about this as well. It's like when you're setting this up, you can actually use Claude Code offline as well with this local setup. So it's really cool.

We'll be testing out today side by side and just comparing how it performs. So the way that you can run this is basically you can just use a quick start here, right? So you copy and paste these instructions directly into your Claude and from there you can start going. Now, if you want this to run with Llama, you can do that too.

And I'll show you how to do it in a second. So let's get straight into this. All right. So you can see how I set it up, right?

If we go to desktop here and basically we open up the desktop, then we pull up this section here. This is actually kind of like an app. It's like a shortcut to run Claude Code directly with our local models, right? Now, if we wanted to use this, you'll see if we zoom in here, right?

You can see that we have Gemma 4 running with this directly, right? So it's local and it's running inside Claude Code. So if we say, okay, are you working? It says, yes, I'm working.

How can I help you, right? Now, if you're using Gemma 4 locally with Claude Code local, it's going to be super fast, but it won't be that powerful. But it'd be good like if you just want to run this, for example, with a lightweight model and you want to run Claude Code for free, then this is a really good option. Again, if you want to use like a bigger model, more powerful model, you got two more options, right?

Number one is Llama 3.3, and the other option is using Quen 3.5 as well, right? Now, this runs locally on a Mac, so it's designed for Mac. You see the three different model options here. And depending on how much RAM you've got, that will change how quickly it responds and how powerful it is, right?

But the main thing is you can run it for free, you can use this locally, and you can run it offline as well. So three big advantages to using this Claude Code local with local AR models, right? And the way that it runs is with something called Google TurboQuant, which basically means that you can run like local models and get more out of them. And you can get the agent harness to work better as well, right?

So we've seen that it works. Let's just test this out now. So if we, for example, wanted to redesign my website, or in fact, let's do something even more basic, right? If we say inside here, okay, create a nice SEO calculator, simple HTML, one page, let's see if it can do that.

The other cool thing is like it has a lot of advantages, you're just using Claude normally, right? So if you're using like Cloud Claude, well, then like, you know, your data isn't local, right? You can't use it on a plane because it doesn't work offline, right? Whereas everything stays on device with something like Claude Code local.

So that's the benefit of this. And you can see it's pretty responsive, it's working nicely. If you want to see how quickly each one works, so Olama, if you're using Olama with this, this is the fastest. Lama CPP is a bit slower.

And then MOX native, sorry, this is faster, right, more tokens per second. If you're using Olama, it's slower. And they've designed it for tool call reliability as well, right? So this is one of the biggest issues with local models, particularly stuff, you know, that's not designed to be agentic.

So local models, normally they struggle with tool calling. But when you're using something like this, it's been fixed, right? So they actually use this tool call option right here to just make it work better with local tools. And you also get the full functionality of using something like Claude Code as well, right?

So for example, if we do like forward slash, by the way, we can ask questions in between, like, by the way, what day is it? And we can ask little side questions whilst it's running on the main process, right? Whilst it's creating the main topic here. And also, this is pretty cool as well.

So like, if you've ever used, for example, like Olama directly with Claude Code, it's okay, but it's not great, right? Whereas if, for example, you use Claude Code with this method, it runs differently, right? It's designed to run in a different process, a different order, so that it actually uses tools better. And it's just a bit more smarter when you're using it.

So this basically means Claude Code free forever. You can run it offline, easy to use. And here's something else that's pretty cool, is you can control it from your phone, right? So it has iMessage that works offline, which is pretty crazy, right?

So if you want to use this with your phone, use iMessage. And then you can use this offline as well on your phone. You've also got this hands-free voice. So if you speak into it, it will reply in your voice as well, which is pretty amazing.

You can control your browser as well. And of course, you can just run, like, Claude Code as you would normally, too. And then you can, if you press Control and O, you can see exactly what it's doing and what it's working on. So that's basically how you can use it, how you can set it up, how you can use it with voice as well.

Pretty cool idea. Really amazing project. Easy to use, easy to set up as well. If you struggled to set up, you can actually just use Claude Code to help you set it up in the first place.

And it seemed to have no problems with that. So, for example, if you go back to the original chat here, you can see that it actually configures it for you and makes sure it's using Jemma 4 inside there. And then it checks it's actually working. So it's pretty cool.

You know, you could use this with smaller models, like 4 billion parameter models, for example. If you're using, like, a M1 Pro, then you could use Jemma 4. If you're using, like, a Max, then you could use Jemma 4 or Quent 3.5. And if you're on Ultra, then you could use, like, all three of the local models that we've talked about today.

So it's pretty impressive. So that's basically how to use it, how to get Claude Code local set up, what it does, why you should care, how it works, etc. We've actually got a 30-day roadmap inside the AI Puffer Boarding, which you can see right here, on how to use it, how to set it up as well, how to install it too. And we have a standard operating procedure as well.

So if you want to get the full guide on this, feel free to get it inside the AI Puffer Boarding. Link in the comments description or just go to the AIPufferBoarding.com. This is my community that's designed to help you save time and grow and scale with AI automation. Inside the community, you can ask questions, get help and support whenever you want.

There's always people online, which means you can get help whenever you need it. Inside the calendar, you can jump on weekly coaching calls where you go deep on stuff like Claude Code, AI agents, Hermes, OpenClaude, that sort of stuff. You can ask questions about your setup, share your screen, etc. You can also meet people in your local city who are doing similar things to you.

So, you know, there's a lot of people inside the community you can meet locally who are using Claude, Claude Code, AI agents. And then inside the classroom, you get access to all my best training, as you can see right here, including new and advanced daily tutorials. So we're always dropping new daily tutorials, as you can see right here, with a new video guide, step-by-step processes and everything you need to really win with AI.

More episodes

Browse all episodes →