AI News Today
← All episodes
Episode 42 · June 20, 2026 · 09:50

Hermes + North Mini Code: New FREE API!

North Mini: The New Free AI Model for Agentic Coding

Discover how to use North Mini, Cohere's new free agentic coding model, with the Hermes agent to build and run tools for zero cost. Learn how to set it up via the Open Router API or locally with Ollama to save your premium tokens for high-level tasks.

00:00 - Intro to North Mini
00:55 - How to Access North Mini
01:18 - Setting up Hermes Agent
01:39 - Running Locally with Ollama
02:45 - Model Benchmarks & Performance
04:07 - Creating AI Agent Profiles
05:21 - Token Management Strategy
07:31 - Building with Free Agents

Full transcript

There's a brand new free API and free local model that you can use with Hermes Agent and it's called North Mini. I'm going to show you exactly how it works. We've run it both locally and also with the API. Pretty powerful stuff, super fast and it's designed for agentic coding.

We've already plugged it into Hermes Agent like you can see and we said like you're working and it actually replies like you can see right here. So it can work with Hermes Agent, it's pretty quick to reply and it can do all sorts of stuff. So if we say for example okay schedule a Japanese practice session at 5am today or 5am tomorrow plus daily as a scheduled task inside Hermes, we can plug that in and then Hermes will use North Mini to just go off and build this. Now you might say at this point okay how does this work?

What are the benchmarks like? How powerful is it? How do you access? I'm going to cover all those questions today.

So number one, how do you get access to it? So you can actually get access on Open Router. So if you type in North Mini you'll see North Mini code is ready to go. It's 256k context window and it's free to use and you can plug it into Hermes Agent directly.

Pretty awesome. Now how do you actually set this up with Hermes? All you do is you go to your terminal like so and then you go to Hermes model, select Open Router from the list and then you would select North Mini code as the model for your Hermes Agent. That's how easy it is.

So just type in Hermes model, set up Open Router. So what is the API? It's this free one. How do you get access to it?

Inside terminal. Now you can also run it with OLAMA. So if you prefer to use OLAMA and run it locally this is not a cloud model from what I've seen. So you can install OLAMA in one single click using this terminal command and then from here you just go to this section and you can see again it's a local model it's not cloud-based which means it's free for anyone to use and run locally and then to pull the model you would click this, paste that into your terminal with OLAMA running.

To run it inside Hermes Agent you can use this terminal command. You can also run it with OpenClaw, Claude Code, CodexApp, Codex and OpenCode as well. Pretty awesome. And so if you are out of tokens or if you've run out of usage on your existing limits you can just use a free model like this and plug into Hermes right nothing can stop you.

That's pretty cool. And so this is how you can use the system how it works etc and this is something I call the free API command engine. So you can wire it into Hermes and you get an agent that can write files, run tools, build things for you for free and you know we just for fun we built this out with Hermes but you can use it agentically and it can call tools. It's actually a very small model so this is a really small model if you're running it locally which means it's really fast and it outperforms Gemma 4 on some of the benchmarks.

It's 30 billion parameters with 3 billion active. It's one command to run it, 256k context window and free whether you use it on the API or if you use it directly with OLAMA as well. So you've got two different options right there and you can see that it's working, it's scheduled the task, pretty simple and easy and you can see that it's currently scheduled that Japanese practice that we just talked about right there. Pretty nice.

So here's the announcement of North Mini and its setup and you can see how it performs on benchmarks. So again you would compare it against models like Quen 3.6, Gemma 4. These are the comparable models if you want to understand how it works etc and you can see the quote from Cohere here who built it. So North Mini Code is Cohere's first agentic coding model.

A 30 billion parameter mixture of experts model with 3 billion active. Optimized for code generation, agentic software engineering and terminal tasks. That is perfect for Hermes Agent. It's also available on HugInFace and OpenRouter for free as well.

So you can see the details on HugInFace right here. In terms of the architecture, how it works, everything else. So how can you wire into Hermes? You just grab a free OpenRouter key, create a Hermes profile for it.

So we actually created North Mini as a profile and you can use this terminal command to create a separate agent profile for North Mini. Why would you do that? I think it's good to separate agent profiles by API because then if one goes down you can use another one. You don't need to switch the models manually and also you can test them side by side and give them the same tasks.

So if you look at our agent operating system for Hermes, you can see that we can select between all these different models and we have the conversation history for each one and we can just use them based on their skills. So we have KimiK 2.7, Quen 3.7 and also North Mini ready to go. We can see the full conversation history over here. We can also talk to our AI agents, we can voice activate them, we can generate images, video and voice with them as well.

But if you just want to chat with North Mini or build out teams with them, you could also use the Kanban board here. So you could actually set up a Kanban board just for North Mini and then give it tasks. It will automatically triage you inside the Kanban board and build it step by step. So now at this point you understand how to set up an agent profile for this, how to run it with free API or with local models.

You've seen it work in action, you've seen how powerful the model is. Like it is a big big update. I mean it's pretty useful. A lot of people run out of tokens quickly.

This is a good option to to fix that. Now let's talk about the old way versus the new way as well. So if you're using an API currently, or every agent uses up tokens, you have to kind of ration those tokens and you're scared of doing stuff because you don't want to use up too many tokens whilst you're building. And that is a big problem my friends, right.

And then also you might want to let it run overnight as a 24-hour AI agent, but if you're running, if you're scared of running out of tokens, well that's a nightmare, right. Whereas if you have a free API, you can run it all day 24-7. There's no meter, especially if it's local. You have a free agentic model wired into Hermes.

You can throw every small job without thinking twice. You can spin up several agents in parallel for free and you can save the premium models for the jobs that actually need them, right. So you can use your frontier models for the really powerful jobs and tools like North Code Mini for the small jobs. Now if you want every system and setup that I've shown you today inside the AgentOS, you can get that inside the AI Profit Boarding.

Link in the comments description, or just go to the AIProfitBoarding.com to get access. And you also get four coaching calls a week, plus daily tutorials as new models drop, a 30-day roadmap, every prompt and the obsidian memory setup for all the agents, and 3,600 members inside the AI Profit Boarding who are building systems like this as well. So you can get that all inside there. Now some people say well free models can't really build anything, but you've seen how it worked today.

Like it can actually be used surgetically. Other people say well setting up a new model, that's a lot of work, but as you've seen it's just like one quick copy and paste terminal command. And other people say well I'll just keep using my paid model for everything, but then you keep paying for the cheap jobs too. Whereas you could route the grind to free models, save the premium for what it's worth, and that's the whole game right now.

And you might say well this sounds technical or difficult to set up etc. We've got 184 pages of testimonials and wins from people setting this sort of stuff up inside the AI Profit Boarding. So if they're non-technical and they can do it, and if I'm non-technical and I can do it, then we can all build with this. You don't need to be a coder or a developer or a technical to build with AI agents anymore.

That has totally changed. So what you just gained? A free coder. Cohere, North, Mini, Code.

Completely free on the API. A real agent because you've wired it into Hermes. It can write files, it can run tools, it can build, it can schedule tasks for you. It's a two-minute setup.

You just set up a profile. I've shown you the terminal commands for that already. It can build. So you have one command and then you get a finished artifact in its workspace.

So for example everything that we build with whatever agent we're using, we get the full setup inside here. So we can see what we've built, we can preview it, and we can come back to it later. So everything that we build with our Hermes agents is saved inside our workspace, which is just awesome and easy to use whenever we want to get there. And then you can run it all day, especially if you're running it local.

So you can have this running 24-7 and doing tasks for you, and you actually have the output because this is an Apache 2.0 Waits license, so you can keep everything that waits. And I would say, you know, honestly the best agents to experiment with are the free ones, because it doesn't matter. Like if you're not going to have to worry about tokens, or limits, or resources, you can just, you know, build cool stuff and see what it does. So if you want to make the free model part of your system, we, you know, there's new free models dropping every few weeks now.

The agent operating system inside the AR Profile Boarding lets you wire each one in and route work to whatever is best and the cheapest without rebuilding anything. So if you want the full agent operating zip file with Hermes in every model inside one dashboard, the setup walkthrough, four weekly coaching calls, 3,600 members, and 155, sorry, 184 pages for members wins, you can get that inside the AI Profit Boarding. Link in the comments description or go to the AIProfitBoarding.com. Inside the community you can ask questions, get help and support in real time.

I answer these questions personally. Inside the classroom you can get access to all my best trainings, including the new daily updates, and we have the agent OS system over here. We update it daily, so we're going to add a new version later today. You get the video tutorials, the zip file, and new tutorials added as they drop.

You can also jump on weekly coaching calls, get help and support in real time. Inside the map you can meet people locally who are building with AI agents like Hermes and North Mini in your local area, and this is all available inside the AI Profit Boarding. Link in the comments description or just go to the AIProfitBoarding.com to get access. Thanks for watching.

More episodes

Browse all episodes →