AI News Today
← All episodes
Episode 104 · August 6, 2026 · 08:00

Maple AI: FREE Local Model on Your iPhone

Full transcript

Today we have a brand new update from Maple Preview. Maple Preview is a new model, open-source, 20b, A1B, ternary weight, reasoning, LLM, state-of-the-art, in its weight class, and it's actually very, very quick to use. So, for example, if we actually go inside our local agent operating system, like you can see, and we say, like, let's say, for example, build out a beautiful snake game, it's actually very, very quick to respond. Look how fast it runs.

Now, bear in mind, like, I've run, for example, Gemma 4 in the past using local models. I've tested all of the latest local models that are actually worth testing on a Mac Studio, and most of them are terrible, or run slow, or slow my whole setup down. This is really, really quick. We can preview everything we've built, we've got everything saved inside our workspace, it's running for free and local, we can switch off the Wi-Fi, continue using it, and it's right up there with Quen 3.6b.

So, let me show you an example of this running side by side. This is the announcement, by the way, and the difference here is that it's really focusing on speed. So, there's a speed quality frontier for local models, and that's really the game that's being played. So, for example, you want quality, but you also want it to run fast.

Usually, the better it is, the slower it gets. So, for example, if you look at Gemma 4, it's right in the middle there. This is the average performance, which means the quality is up, and this is the speed. So, Gemma 4, E4b, E2b, both in the middle.

If you look at, for example, LFM 2.5, pretty fast model, actually, when you use it, that runs fine. Now, if we look at Maple Preview, according to their benchmarks, now, I've tested it myself, it is pretty decent. It's not the best in the world at coding, I'll tell you that for free, but it is pretty fast when you're running a local model. It's average performance is way up there, and its speed is up there.

Now, this is what it's all about, really. So, if you look at models in the same caliber, like, for example, Quen 3.6 27b, which is probably the one that I hear about the most, but it's way too slow to run on my setup, for example, like a Mac Studio, Quen 3.6 27b is a lot slower. The other option is 1-bit Bonsai 27b as well. That's considered a decent model, which is basically Quen 3.6 27b.

That could take like five minutes to respond to a simple question, depending on where you run it. So, if you look, for example, at Maple Preview, they're looking at running it on an iPhone. That's the other legendary bit about this, is like, it's focused on smaller devices, it's focused on mobile, and I think this is a future that's coming where we can run local models for free on a mobile device, which right now, if you look at the comparison, the speed is not really possible with something like binary Bonsai 27b, aka Quen 3.6 27b. So, it's designed to be very light, but at the same time, run pretty nice.

Now, here's another example. So, this is running on a MacBook Pro, where it can be a lot more autonomous. So, they've said it autonomously decides to remember details, and that's the other big advantage here, is that it can adapt, it can improve, it can learn. Now, they've actually compared it against Claude Sonic 5 with the same scenario.

I mean, when I've tested it on coding tests, it's nowhere near the same level as Sonic 5, so I'm not even gonna entertain that idea, my friend, but you can see where it goes on the performance benchmarks. Now, you can test it out on their website if you just want a quick test. It's available at chat.deepgray, or you can get the open weights directly on Hugging Face. Now, I've already plugged it into my Gentic operating system, which you can see over here, and we have the local section.

The other great thing about that is we can switch in and switch out any model that we want to, and basically code for free. Also, I quite like the way that it responds. So, it feels quite nice when you're using it. It gives you some detailed responses.

It certainly gives better responses than a lot of the other models that I've tested. Let's say, for example, build out a beautiful landing page for an SEO agency, and then we can leave that running over here. It'll run locally for free. And also, the great thing about this is that it codes really fast, but we can also use our other agents inside the OS whilst that's running in the background.

So, for example, you could have LFM, which is another model I was testing earlier today, and that works really nicely with Hermes Agent, and then you can power your whole agentic operating system using free models. You've got the local builder over here, and you have LFM over here. Now, you might also wonder, okay, what does ternary weight reasoning mean when it comes to local models? This is something I actually learned today.

So, normal AR models store every connection as a precise number with many decimal places. Maple stores each one as just minus, zero, or plus. So, just three symbols. That's what ternary weights means.

So, writing the brain with three symbols makes the file tiny and the mass fast, which means a model that would normally need around like 38 gigabytes can fit into five gigabytes, and your max chips, for example, can run through it way faster. Now, also something to note here is that the second trick is the expert system. So, Maple holds 256 small specialist networks inside it, and for each word it generates, a router wakes up only the eight most useful ones. So, you get the knowledge of like a 20 billion parameter model, in theory, with the running cost of a one billion one.

And then you get a 128K token context window and a MIT license with it, which means that it's free, you can use it commercially, and you get an actual reasoning model. You also might think, okay, if you squashed the numbers, does it give the same quality of outputs? So, when I've tested it myself on the outputs, I wouldn't say it's as good as something like Gemma 4 when it comes to actually coding, but it runs way faster. So, I mean, you can see the benchmarks here and how it performs.

For me personally, did I get those sorts of results? No, like it wasn't as good at coding as something like GLM 4.7 Flash or GPT-OSS when I've tested it. But test it out for yourself, see what you think. The main benefit that I think here is that you can use it on mobile devices, and that's a massive win.

So, it's easy to set up, it's fast, it's free, it's local, it can code, it's got an interesting architecture, and it plugs right into an agentic operating system. So, if you want to get our full system with the local engine here, and you can swap in and swap out any sort of models you want to use, you've got, for example, the memory system that plugs into all of this, and you can customize this as much as you want. So, it's got a full mission control with all your favorite agents here. So, for example, we actually tested LFM 2.5, 2.6b, that dropped to date as well with Hermes Agent, very powerful model.

And if you want to get this whole setup, you can get it inside the AI Profit Boardroom community, link in the comments description, or go to the AIProfitBoard.com. And inside this community, you can ask questions, you can get help and support in real time. I personally answer these questions each day with a video tutorial. Inside the classroom, you can get access to all my best trainings, new lessons, and grab the full agentic operating system with free local models here.

Also, we have a full training on how to run an agentic operating system for free too. And then you can jump on weekly coaching course, ask questions, get help and support in real time. And inside the map, you can meet people locally near you who are building with AI agents like you. So, feel free to get this, link in the comments description, or just go to the AIProfitBoardroom.com.

Thanks for watching.

More episodes

Browse all episodes →