Discover North Mini Code, a revolutionary 30B parameter model that outperforms models four times its size while running privately on your machine. Learn how to set up this agentic coding engine with voice control and Ollama for a seamless, offline developer experience.
00:00 - Intro to North Mini Code
00:16 - Benchmarks vs Larger Models
00:53 - Live Demo: Building by Voice
02:54 - MoE Architecture Explained
04:40 - Key Specs & Context Window
07:00 - Agent OS Integration
07:37 - 4-Step Installation Guide
Full transcript
There is a brand new local model called NorthMiniCode that is super powerful for coding and you can also get access to it for free. So this is available on Olama now. So NorthMiniCode and Olama are working together and it's pretty powerful as you can see right here. So if we look at the benchmarks it's a little 30 billion parameter model that scores models four times its size according to the artificial analysis coding index.
So you can see for example how NorthMiniCode which is performing at 33.4 here is outperforming DevStraw, Mistraw 4 and Nemetron 3 Super which is unbelievable. Now we've got it running inside our agent operating system already so if we go to the local section here we can build with it and we have NorthMiniCode plugged in. You can also plug this into your AI agents and that means it's free, it's local, it's private, it can work offline and it can actually build stuff. So for example let's give it a little test run here just as an example.
Build a colorful to-do list app and then we can control this with our voice, we can start building with it and you can see that's working right here and so this is local, it's ready to go, it can actually build stuff and it's free to use which is unbelievable. Now if we actually go to the preview section here we can see what we've previously built so everything that we've built is saved and then you can see that is quickly coding out here. Now once that html is fully coded out we can actually preview it as you can see and if we test it it actually works, unbelievable right. So this is a free local model just dropped, works properly, works offline, works with Hermes agent or whatever agent you want to plug into and also we've got a workspace here where you can use your local models with this whole system and it's really good and ready to go.
So you might be wondering okay how does North Mini Code work, how does it perform etc. Let's pull up the benchmarks right here. So if we compare it on benchmarks it's not quite up there with Quen 3.6 but you need a good setup for Quen 3.6. However it is outperforming Gemma 4 as you can see on Terminal Bench, it's also outperforming Gemma 4 on Terminal Bench Hard, SW Bench Verified, it's absolutely crushing Gemma 4 and it's right up there with Dev Strahl Small 2 and it's up there with Quen 3.6 as well surprisingly.
So it's an agentic coding model that's pretty powerful and easy to build with as you've seen. You can even control your voice, you can build stuff with it, it actually works and it's ready to go. So pretty powerful stuff. So let's talk about the local AI coding engine.
This is a free coding model that lives inside your computer, it's fast, it's private, it's made for building and five things make it work. So number one it's a coder, so North Mini Code is trained for one job which is writing software. Number two it's a small body so it's only 30 billion in size but only a few experts actually fire up a word so it's quick to run as you saw and it outscores models four times larger at coding which is unbelievable. So you can outperform models four times its size because it's small but it's a mixture of experts model.
It has reasoning on so you can switch on reasoning whenever you want to, it's free and it's under the Apache 2.0 license, works offline too and you can build by voice as well. So let's compare these side by side. You know you could pay for a subscription but that would get expensive or you can use a tool like this to basically code unlimited. You've got a free coder built only for software, it beats models four times its size, it works at 92 words a second on a normal mac with no limits, nothing you type ever leaves your machine, you say the word and the whole app appears as you've seen today and the result is a fast private coding engine that you actually own.
That is unbelievable. Now how does it work? Why is it so quick? So it only wakes a few experts per word.
This is what makes it so fast. You can run this on a laptop too. So old models wake the whole brain for every single word which means big models are quite slow locally. North is split into 128 experts and only eight of them fire for each word.
So it's a giant brain with tiny effort and that's how a 30 billion parameter model runs as quick as a small one as you've seen. So what's actually inside it? Well it's 30 billion parameters total size but only three billion fires per word. So eight of its 128 experts, that's where the speed comes from.
256k context window which is pretty good. That is pretty impressive. Now bear in mind like Kimi K2.7 which is a really powerful model, that only has a 256k context window. So it is good.
Then you've got three to one attention mix which means it's quick on long files and also is RLVR, trained on real work with reinforcement learning with verifiable rewards and real software and terminal tasks, not just text. It has native tool use so it's built to drive coding agents, call tools, run a terminal, think between steps. That is great if you want to plug it into Hermes agent and on Apache 2.0 it's free and open right. You can use it for anything including commercial stuff which is pretty nice.
And also it was trained across real coding agents, so sw-agent, open code, terminus and more. So it's good at messy real world work, not just clean benchmark puzzles. Now you might wonder, okay how big is the file? So it's 19 gigabytes at Q4, so it fits a 16 to 36 gigabyte Mac with room to spare.
92 words a second on my M4 Max which I measured, so it's quicker than the model's half its size. It punches above its weight so it outscores Nemetron 3 Super, Mistrial 4 and Devstrial 2 and it scored 33.4 on coding index which means that it beats Quen 3.5, Gemma 4 and Devstrial 2. Now I actually tested this just on a quick snake game test. If you don't have the thinking on, it takes about 20 seconds to build.
If you have thinking on, then it takes about 48 seconds to build. And you can see some examples of just, you know, interesting fun stuff that we built here. Now we actually keep it warm to make it feel instant as private as well, so we keep it warm in memory so it never has to reload. And that just means that there's no lag when we're using it.
And then also it runs on a loop inside our Mac, which we actually checked as you can see inside the terminal here. Now if you want North and this full coding engine inside your dashboard, you can check it out. We've actually already built it into the agent operating system, so you can get that inside the AgentOS with the voice input, build to previews live, every app is saved in the workspace and it sits next to my cloud agents as a free fast coder. Plus we've already plugged it into Hermes agent as well, so if we go to Hermes here, we can use this model with Hermes agent too, which is unreal.
We've got the local agent ready to go right there. So you can run it in about 10 minutes. You just need to update olama, pull North in, keep it warm, and then wake it once, right. And if you're wondering how do you do that inside your terminal, here are the instructions you can use to run this inside your terminal quickly, step by step.
So it's done in like four steps as you can see. And it's pretty powerful. Now some people say a free local model can't really code. I mean it outscores models four times its size, so that's pretty impressive in itself and you've seen it actually build stuff today.
You might also think well a 30 gigabyte class model will be way too slow. Only fires like eight of its 28 experts, so it runs at 90 words a second. It's pretty good from what I've tested so far and you've seen it work live today. Also you might say well setting up a local model is too technical.
Literally it's three short terminal commands. So we give you the commands, you can see them right here. This is how you can run NorthCode mini on your machine. So just to recap, here's what you learned today.
You learned how to run a free model on Apache 2.0 that's a free coder. You learned how it actually outperforms Giants four times its size, which is unbelievable. Again I would test this all for yourself. Don't like believe all the benchmarks.
Test out yourself. See what you think. If you like it, great. If you don't, try a different model.
You learned how to stop waiting because it's pretty fast. You learned how to run it privately and locally and you learned how to pick between speed or depth. So you can switch on thinking on or off and you can build by voice as you've seen today. Now if you want to get the full system from me, you can get that inside the AgentOS system and this is inside the AirPuffer boardroom.
So you get the full AgentOS with North, Cloud, GLM, Hermes and more inside one dashboard. You get the local AI coding engine setup with the voice control, the previews and the system where it just runs warm ready to go. You got a 30-day roadmap inside the AirPuffer boardroom and you can get this all inside our community here. So if you go inside the classroom once you've joined, link in the comments description and go to the airpufferboardroom.com.
Then go to the AgentOS system here. You can grab the video tutorial. We've got the last update date on this host. Set up the resources with the zip file and we add new daily tutorials as you can see to get the most out of all this stuff.
Inside the community you can ask questions and each day I answer them with a video tutorial so that you can get as much help as you want. Inside the calendar you can jump on weekly code duels, share your screen, ask questions in real time and then also you get a map where you can connect with people locally who are building with AI agents just like you've seen today. So if you want to get that all inside the AirPuffer boardroom, link in the comments description or go to the airpufferboardroom.com. Thanks for watching!
More episodes