AI News Today
← All episodes
Episode 56 · June 30, 2026 · 15:11

FREE Agent OS Engine: Run Agents For $0

Full transcript

Asian operating systems are super powerful. I mean you can basically build and automate any sort of workflow that you want inside a setup like this. So for example we've even built like a outreach tool with Hermes Asian. We've got a voice activated version of Hermes over here.

We have a memory galaxy that's ready to go. We can automate videos in like one single click. We can also for example automate and create SEO content and deploy it to our websites in one single click and this is all inside one beautiful system that works together to orchestrate and manage our agents. However one thing that I see people worried about all the time is like how can you run these agent operating systems for free or get the most out of them and that's exactly what we're going to talk about today with the free Asian operating system engine so you can run a whole operating system of AI agents for free forever using a combination of all the different systems I'm going to show you today and the thing I would say here is like you know a lot of people worried about token usage, they're worried about APIs, they're worried about you know having a system like this but then you know going through their limits on the CLI or their coding plan for the month and so the whole point of the Asian operating system is of course to let it run and automate autonomously but if you're worried about doing that then let me show you some ways around that so you can use free models to help you as much as you can.

So let's get straight into this and there's basically five systems I'm going to show you today that will help you either completely run this for free or reduce the amount of tokens used massively so we're going to break down each one of these as you can see so free local models, free APIs, CLIs, token optimization and also free memory systems so you can plug this into your agent operating system and then you can have these amazing automations and workflows running but you can run them for free so let's get straight into this. The first thing that I would say is that you can run local models and I've tested out quite a few recently I'll show you what I think is good and what I think is bad right and I only have like a I have a Mac studio even then it's not amazing for running local models but the best one that I've seen so far is the Quen 3.5 27b coder that's been a pretty good one that we've used so far I can create some awesome stuff let me show you an example of something that we built here so this is actually fully built with a local AI agent we also for example created this system which is is not like I mean it's not frontier level right it's not going to be like opus 4.8 but if you want to run this for free and you you know for most things you're not going to be creating like crazy graphics or video games you're just going to be creating a lot of content and that sort of thing so if you want to know how to use these well you can use a system like I'm showing you right now and so we have the local engine here and we can plug in free models into this system free local models and then we can build and automate whatever we want so for example that game that just showed you a second ago that was fully created with this system here and it's the same with free cloud code you can run local models and you can run free apis for free cloud code which is an open source project that basically takes your free models and then plugs them into the agent harness which is free cloud code and so that is method number one which is using free local models now if you want to indicate where do you get free local models from you can get them from Hugging Face and also Olama if you're running local models on a Mac then you can also use MLX to run them as well and that seems to be better than Olama and I've tested them out personally they've been a lot faster for example so that is method number one and if you want like a local model to actually be useful you can check out my local model benchmarks that I've tested pretty much everything with so you can see we've run 42 different tasks for each of these agents or NIF is another good one as well for coding locally and so if you want to see what we've built and how they compare you can check out GoldieBench for that and we've got a local leaderboard for this sort of stuff next up you can actually use and wire in your CLIs that you're already paying for for free into this system right and also what you can do here is you can use anything that you're already using inside your agent OS what do I mean by that well for example let's say recently I subscribed to glm 5.2 if I've got that as a CLI then I can also use that inside my agent operating system I would just plug in the CLI into the agent operating system and then it doesn't cost me anything extra and the good thing about for example the coding plan on these open source models like glm 5.2 or kimi k217 the great thing about them is that they don't tend to run out of tokens right you can use them a lot get a lot out of them but you don't need to worry about that so if you've got existing CLIs like for example Claude as well you can plug that into your agent operating system too and the great thing about that is you still get the power of an agent operating system but you don't need to pay anything extra so that is method number two which is also free now what you can also do from there is you can use free APIs so as an example of that if you go to open router here and you type in free you'll see loads of different free APIs that you can get access to and you can plug into your agent operating system so for example north mini code is a free API that we can get from open router we can grab a free API key from open router then we can go back to Hermes agent for example and we can plug that free API key directly into Hermes agent so we've actually got a separate profile for our free API key with north mini code and we can use Hermes agent for free whenever we want and so all of these workflows for example like the Hermes oracle or the outreach engine that we've built for Hermes agent you can actually plug a free API into that and use them that way so that is method number three and just to recap here we've talked about using CLIs we've talked about using local models and also free API keys as well also the great thing about this is for example if it's a basic task you could route that to a free local model if it's like a medium difficulty task you could give that to a free API and then when it's a really hard task or something that requires a front-end model which is not that often then you can plug it into a CLI that you already own for example like Claude or GLM 512 but either way you're reducing the amount of tokens you use you're reducing the amount of resources required to run an agent operating system and it's also super easy as well so I'll give an example like our Claude plan that we use a lot we actually built the agent operating system using Claude and that comes with a CLI and so you can plug your CLI into the agent OS it can connect to the whole system the same way for example if you're if you're already subscribed to Twitter you can plug in grok build and you can plug that into Hermes agent you can plug that into the grok build section over here and then just by having a Twitter subscription that we already have I can plug that into my system and build and automate anything that I want and you know grok itself is pretty good like you can you can generate some pretty cool stuff as you can see here so it's really about just being resourceful and finding ways around things if you're worried about this stuff you're not worried about it that you don't need to think about it but a lot of people do ask me this question so that's why I'm showing you today now also what you can do is you can make your existing AI agents way more efficient so what I mean by that is basically sometimes you can be using too many tokens for smaller tasks without even realizing it so what you can actually do from here is you can use a token optimization layer like for example headroom that can sit in front of the jobs that you give to your AI agent now if you're wondering how to access the headroom you can see an example right here it's actually designed to cut your token costs by like 50% here's the main github it's actually trending on github on the research it's designed to use like 60 to 95% fewer tokens on the same answers so what that means is basically if you give headroom a task a task that you use with headroom might take up 60 to 95% less tokens than if you use a normal sort CLI so it's kind of like compressing the amount of tokens you use with this compression layer now when I tested it myself I think it reduced it by like 20 to 30 so it depends on the task you give to headroom and what you're working on which CLI you use as well and also what your memory system is like but that's just another way to use this and you might say well isn't shrinking the context going to make the answers worse sometimes it can actually make them better because a bloated prompt buries the actual task under noise the model has to wade through so if you trim to what matters that means the model spends its attention on your real task not on rereading the same context that you have for the hundredth time and so you send less and then you use less and the answer is actually sharper so that's a really good example now the final step and we've talked about four different ways to run an agent operating system for free and reduce the amount of tokens you use so number one was free local models number two was using free apis number three was using CLIs and number four was using token optimization systems like headroom the final system is free memory right so you can actually use a free memory system so for example we have obsidian plugged into this system and let's say for example we're using any of our ai agents whether that's claude whether that's for example hermes open claw whatever we can use free app like obsidian as the memory layer for our ai agents and that means they all have context and it means that they can also organize your context for you and so if i ask hermes a question about my goals or my systems my business it knows everything that it needs to know about me in an efficient way i've never had to organize this myself but it's it's a pretty amazing layer that i can just plug into any agent if a new agent comes out tomorrow no problem i can plug in obsidian directly into it right and obsidian is free to use just a simple app like this literally all it is is just like a bunch of markdown files organized into folders and then organized properly right and every single one of these dots is a dot that links to me and my system my knowledge graph that my agents can read and also the great thing about this is it's a positive feedback loop so your agents update and organize your notes for you every time you use them so if we go into the recent section here you can see these are all memories from when i've used my agent and it's updated every few hours as you can see and then also it's a positive feedback loop because when i'm using my agents i don't need to re-prompt them i don't need to paste in context or correct them because they pull out the context from this system and so they update and organize it for me which saves me time and then it also saves me time on the other side because when i'm using my agents like hermes we've got this free system that automatically updates them on everything i've been working on up to the hour right so if i'm if i go into for example hermes or claude i'm like what did i work on yesterday or what did i work on a few hours ago it can pull in this context and all these memories that we've used to get the most out of the system and that's how powerful it is so that is method number five for running an agent operating system for free right which is the free memory system and you have your obsidian vault with your goals your clients your voice and that just organizes itself and you can use free apis or free clis plugged into your ai agents to update and organize it for you it's a free memory loop so that's basically the whole system and if you look at like the old way versus the new way for using this you know the old way is like every agent is using too many tokens because you don't have hebron plugged in you might be using loads of different subscriptions but then also paying for the api without even realizing it you're worried about using your agents who use them less you have fewer agents you're worried about the prompts and also if you're not using local models and let's say for example you have data that's private well that would go to the cloud as well and then sometimes you hit a rate limit and then you have to stop mid build and the result with the old way is that you've got an agent operating system you're just not interested in using because you're worried about too many tokens with the new way you've got a system where you can run it on free local models which are private offline they can work without the internet so your agent operating system can actually run offline you can plug free apis into your agents as i showed you before you can get them from open router you can use frontier level clis like clort and plug those into your system you can use token optimization systems like headroom to reduce the amount of tokens you use and you've got a free obsidian memory which means your agents don't burn tokens relearning your business and the result is you can let it all loop and it's free to use so this is pretty powerful stuff now if you want everything that we've talked about today and just to recap you stopped and reduced your token usage by 90% with headroom you got free cloud models with apis you learned how to plug your cli's into this system you learned how to make your agents more efficient you got a free memory system and you can finally let it loop and run without worrying about all this stuff you can get that inside the agent operating system right so we have this inside the air profit boardroom and if you want the free agent os engine ready to go you can build every pc yourself but if you'd rather skip it and just run my exact setup it's all done for you inside the ai profitable so you get the full agent operating system you can you get all our trainings on the best free apis you get access to me so you can ask me anything and anytime you need we've got token optimization playbooks inside there you get the full zip file on every prompt for building the agent os ready to go and we've got loads of members who are using systems like this to get the most out of it all right so you can see we've got 194 pages of testimonials from people inside our community getting the most out of this stuff so if you want to join us and be part of this feel free to join but the main thing i would say is we're all learning and growing together and just building amazing stuff as you can see right here it's also super exciting to build an agent operating system so if you want to get that you can get it inside the air profitable community just go to the classroom and then go to new daily updates and you can find the agent os here with the video tutorial the last update date the guide and also a zip file on how to install it we had new daily guides based on what's actually useful so if you're interested in local models we actually did a new test with a full step-by-step guide here on quable and if you want to ask questions i answer them personally inside the community and so does the rest of the community to help you you can also jump on weekly coaching calls ask questions get help in real time etc and inside the map you can meet people in your local area who are building with ai agents like you and you can ask me any questions anytime you want to so thanks for watching hope to see inside there link in the comments description or go to the aiprofitborn.com to get access

More episodes

Browse all episodes →