AI News Today
← All episodes
Episode 67 · July 7, 2026 · 08:52

How to Run Hermes Agent FREE Forever!

Run Hermes Agent FREE Forever with OmniRoute (231+ Providers, Auto-Fallback & Token Compression)

The script demonstrates how to run Hermes Agent for free using the newly released open-source OmniRoute/Omniroot gateway, which routes requests across 231–237 providers with millisecond auto-fallback when a free API is rate-limited. It explains that OmniRoute compresses tokens using RTK (removes repetition/duplication) and Caveman (short, blunt outputs) to reduce rate limits, and offers one endpoint that Hermes can use without β€œknowing the difference.” The creator shows testing Hermes with an OmniRoute profile, building a landing page locally, and recommends using separate Hermes agent profiles per API/model to compare side by side. Setup is described as non-technical via simple terminal commands and GitHub instructions, with examples of free-forever providers and optional OpenRouter integration. The video also promotes the AI Profit Boardroom and its Hermes Agent OS workflows, memory system, community support, and coaching calls.

00:00 Free Hermes Setup
01:01 OmniRoute Benefits
01:43 Hermes Profiles Demo
02:57 How Routing Works
03:55 Workflows and Memory
04:33 Dashboard Setup Steps
05:51 Free Providers List
06:29 Token Compression Explained
07:01 Wrap Up and Results
07:23 Agent OS Offer
08:24 Community and Coaching
08:50 Final Thanks

Full transcript

Today I'm going to show you how to run Hermes Agent for free forever using something called Omniroot that just dropped. It's a brand new free way to code across like 231 free providers and so what you can do with Omniroot is you can plug this into like pretty much anything like Clawcode, Codex, I've already done tutorials on that. Today we've tested out with Hermes Agent and I'll show you exactly how this works step by step today. So basically with Omniroot this is a way to use coding tools for free right loads of different free APIs inside there.

Now what you can do with that is because it's open source you can plug that into Hermes Agent so you can see for example we have Omniroot which is the free profile with Hermes Agent and so for example we test out here yes it works fine okay so for example we say okay I don't know um code out a beautiful landing page we can use that like so and if you're wondering okay how does this work so essentially this is an open source project and it means that it's got auto fallback across 237 providers and that was in milliseconds so like if one of these free APIs gets rate limited or you can't use it no problem it automatically switches to the next one. It's also got RTK and Caveman built in these are both ways of compressing the number of tokens you use so you reduce the amount of tokens which is great because if you're using free APIs then quite often you can get rate limited and this reduces the chance of that happening. Also it's free to start there's 90 providers with a free tier and 11 of them are actually free forever which is pretty cool every tool actually works there's one endpoint and it's production grade and so if we go back into Hermes now it's beginning to use that system with Omniroot. The way that I like to set this up is I'll have a different agent profile for each Hermes agent I use based on the API and that means I can just test them out side by side I have different conversations and different conversation history for each I can use two Hermes agents at the same time and every time I want to test something like this we can go directly into it.

So for example here you can see that Hermes with Omniroot is selected but for example like Tencent HY3 just dropped today we already plugged that into the system as well. If you're wondering like how do you actually use Omniroot with Hermes agent I'll show you that in a second how to set up. So basically the way this works is you got Hermes agent which is the hands then you got Omniroot which is the brain and then that plugs into Hermes to actually get stuff done. So for example here you can see that we've got the website fully built with Omniroot so if we just run this terminal command like so we can open up the page we just built with Hermes and Omniroot pretty nice pretty clean.

So you can actually code out locally you can do agentic tasks as well using Omniroot and you can use it for free forever with Hermes agent which is pretty mind-blowing in itself. So let's talk about how this works step by step and how to use it. So there's three things that happen automatically. Number one is compression so reduce amount of tokens you have with RTK and caveman.

Number two routes across 93 providers and number three you can switch in one word between those different models. So Omniroot is basically like a small program that runs on your computer it's a gateway like a post office that knows how to reach every AI provider in the world. So you wire Hermes to that gateway instead of talking to OpenAI or Anthropic directly and Hermes sends a request Omniroot reads it compresses it picks the right provider and routes it. So the model answers and Omniroot sends the answer back and Hermes doesn't really know the difference.

So this is pretty much how it works. Now the cool thing about this is you can prompt it then you can use the Hermes profile for Omniroot. Omniroot compresses it it selects a provider automatically all this is happening in the background automatically without you doing anything. So what you're doing is just chatting to it like you would normally with Hermes but the difference is that it runs through this whole model right here.

The other cool thing is you could plug it into custom workflows. So for example we have Hermes agent with a custom workflow for lead generation and sending out emails and managing our inbox and we could easily just plug the model for Omniroot into this system give it access to our emails and then it can go off and send emails as well really which is pretty cool. Now also what we have over here is a memory system. So with Omniroot we can plug in Hermes agent into our memory system and then it can recall context and knows everything about us straight away.

So as soon as we switch the API no problem we've also switched the memory and everything else because it works as one powerful system inside this full Hermes agent OS system that we've created. Now you might say okay how does this work when it comes to actually setting up Hermes with this model. So the way that we did it if we have a look at the profile inside our dashboard. So if we go to manage inside our dashboard here and we'll wait for that to load in the background.

So we can go to our profiles then we're going to find Omniroot which you can see over here and then that's routed to Omniroot directly. So this runs via this gateway and then we set it up inside our profile. The cool thing about this is like the old ways you have one provider one API key one model it uses up a lot of tokens the provider goes down sometimes which means Hermes stops and the agent stops as well and then you only have a few models to choose from. With this new setup you've got 237 providers one endpoint locally rtk and caveman compressor tokens so you don't use too much with your agents it switches the model automatically Omniroot is local so never goes down never limits you either and then you've got three models running in the background with it directly.

Now if you're wondering how to get Omniroot running it is non-technical so some people say you know you have to be technical to set up. No no no so you just run this command and then you run this right and if you want like the full details you can just go to their github to set it up. Now you may also say these models are not free but they actually are. I saw like some people mention it on my tutorial the other day but no it's free and we've tested it and it actually works and there's a bunch of free forever providers as well so you've got open code zen, pollinations, cuoda-a and cairo-ai2 so these are all forever free options.

You can also add open router as well so if you want like 353 more models behind one key then you can set up open router 2 and that has a free router 2 and then if you want to set up an api key for hermes you can just use a terminal command like this. Now you might also wonder okay how does the token suppression work like because that's pretty useful for whatever different api you use. So rtk basically reduces like repeated patterns and the duplication of fluff and caveman reduces the fluff on the way out so what caveman does is it makes your ai speak like a caveman on the output tokens so when it replies to you it replies very blunt and very brief which reduces the amount of tokens your models use. This is also super useful if you're using like for example claudor something like that so that's basically it that's the whole system we've shown you how to use omniroot, how to use it for free forever with hermes, how it actually works, what it can build, what the output is here.

I wouldn't say this is like fable 5 level but at the same time it can build stuff. I've shown you what it built it works pretty nicely it's pretty smooth it runs for free and there's a lot of free providers inside there plus it's pretty easy to set up. So if you want to get our full agent operating system we've got loads of great workflows for hermes so if we have a look here for example you've got the talk mode, you've got the voice agent, we have hermes oracle for tracking the latest news, we have hermes astros that looks at your competitors and then gives you content ideas with new angles, we have the outreach store, we even have a mixture of agents plugged into the system and we've got a memory system here to give all of our agents context in like one click and we have loads of cool workflows for seo, for video agents, for music, for loop engineering, whatever you want. So if you want to get that full system it's inside the aiprofferboarding link in the comments inscription or just go to the aiprofferboarding.com and if you want the full setup for me you can go to the classroom then go to new daily updates and you get the agent os here.

You could build this out yourself I spend about three or four hours a day building this out and improving it so if you want my setup which I think it's a lot easier and faster to implement you can get it right here we have a video tutorial you can see when it was last updated you can get a full guide on setting up plus some resources and we update this with new features all the time so for example omni route is a new system we just set up into the agent s as well inside the community you can ask questions get help and support I personally answer each of these with a video tutorial plus you can get access to all of the other um community here and there's always people online 24 7. Inside the calendar you can drop a weekly coaching calls get help and support on time inside the map you can meet people in your local area who are building with ai agents like you and that's all inside the aiprofferboarding link in the comments inscription or just go to the aiprofferboarding.com thanks for watching

More episodes

Browse all episodes →