Full transcript
K3 from China is absolutely amazing. I want to show you some examples of what we've built with it, but so far, let's recap on what we've got. So Kimi K3 just dropped yesterday, about 24 hours ago, 2.8 trillion parameter model, 1 million context window, really powerful, open source, comes out from China. And also you can use it on Kimi.com, Kimi Work, Kimi Code, Kimi API.
And it is right up there in terms of performing with Kimi versus Fable 5 versus GPT 5.6. So, and I'll show you some comparisons in a second, but essentially you can see that it's doing very, very well on the benchmarks. This is not just like an open source model. This is not like GLM 5.2 where it came out.
It's like, oh yes, it's kind of right up there with Opus 4.8. No, no, no. This is Kimi K3 on certain benchmarks outperforming GPT 5.6 Sol and Fable 5. So this is not hype.
I'm going to show you exactly what we're testing out as well with this. So on GoldieBench, we tested out 50 different tasks. And if you check out the quality of this stuff, like it can create really nice 3D games. It's got a nice ambience and feel to everything that we've created.
Here's like a full open world game that we created with Kimi K3. Just looks absolutely awesome. We actually generated like a kind of copy of the style of Skyrim. So we use it for inspiration.
And you can see if we refresh the page here, like even the opening video looks beautiful, like this is fully custom made from K3, less than 24 hours ago. So it didn't take a long time to create, but the outputs of this stuff is unbelievable, just looks beautiful. And this is like a huge open world game where, you know, the details, the ambience, the vibe of it just looks awesome. I can't see any bugs in it so far as well.
It just feels really smooth when we look around. It looks fantastic. And we generated this in like, what, one single prompt. So it's unbelievable what you can do with it.
Now you might be wondering, okay, how does it perform versus something like Fable 5 and GPT 5.6 in real benchmarks. Let me show you what we tested so far. So we ran a bunch of tests of GPT 5.5, sorry, GPT 5.6 versus Fable 5 versus Kimi K3. And if we have a look, for example, this is the output for like a Minecraft style game from Kimi K3.
Super nice, works perfectly. It's even got water there and the graphics work nicely. If we compare that versus Fable 5 and GPT 5.6, this is the output from GPT 5.6. Sorry, Fable 5.
It looks okay, but it's a little bit buggy in places. And then also the controls are not that nice and it doesn't have that water vibe that we had a second ago. And then GPT 5.6 came nowhere near. Look at this output.
It just doesn't look anywhere near as good. Let's have a look at a Dragon Realm game that we created. So I've already shown you this example from Kimi K3. If we have a look at Fable 5, it's not bad, but it just doesn't look as good.
It doesn't feel as good. There's not as much detail, et cetera. And the same with the ambience inside GPT 5.6. This actually feels like almost like an AI generational gap between K3 and GPT 5.6 when we compare them, like the graphics, the detail, the vibe, the colors, nothing is quite as good as Kimi K3 when we're building with it.
Now for me personally, like I like to use all of these models and I'll have them side-by-side running inside our agent operating system, which you can get link in the comments, description, or go to the airprofitable.com. And with this system, you know, you've got Kimi Code, you've got, for example, Claude, you have GPT 5.6 all working alongside each other. However, what we've also done is plugged in Kimi K3 into Hermes agent and it runs tool holes pretty nicely as well. So it responds pretty quickly when you're using it and you can use the coding plan to plug this into Hermes agent to use it agentically.
And then for example, we tested it with a tool hole like forward slash learn, and then plugged in a guide here, as you can see, and actually read the page, detailed all the lessons that it learned from that particular topic, and then actually gave us a link to the skill that it created so that it can remember how to use that topic forever. And it works really, really smoothly. Something else that we did is we actually linked the Blender MCP over to Hermes. And Blender is like a really technical tool for 3D modeling.
It's great for like product shoots. It's great for coming up with videos, et cetera. But the problem is that it's really technical to use. So you can actually get Hermes agent to operate it.
If you have a good model plugged into it. Now, one of the few models that I think you can actually use something as technical as Blender MCP with is Kimi K3. And you might say, well, I don't want to use Blender, but the point here is like it can actually use any other sort of software. I mean, this could be for NA10.
It could be for WordPress. It could be for whatever app you want to use. The point is you don't really need the skill yourself anymore because you've got Hermes agent as an AI employee that can operate the skills for you along with Kimi K3 and it works really nicely. And I think this is really the first time we have ever seen an open source model catch up or potentially overtake something as frontier as Cable 5 or GPT 5.6.
It is really at the point now where like China, number one, they're releasing models faster. So Kimi K2.7 only just came out last month. And also bear in mind, this is cheaper as well. So if you're using the API with Kimi K2.7 or Kimi K3, it's a lot cheaper than using these frontier models from the US.
And it also self-improves itself, which I think, you know, some people are common to say, well, every model self-improves itself. But the thing that I've seen is that Kimi K3 is just rapidly improving. So I wouldn't be surprised if you see like Kimi K3.1 come out next month or Kimi K3.3 or something like that. And it will be even better than what we're looking at.
But the difference is that if you look at the way the US is going with its models, like they're always going to be in preview before they get released. So Fable 5 came out three days later, got taken down. We have to wait about two weeks and then it came back out again, but it's not going to be on the subscription forever. Very, very complicated system.
And then you've got, for example, like GPT 5.6 came out in preview. We all have to wait two weeks. It has to be tested by security firms before it could get released to the public. And even now, like when you're using GPT 5.6, it runs out of tokens very quickly.
With Kimi, they're releasing new updates faster, more openly, making everything open source, and then also it's way cheaper to use. Like for example, if you actually look at Kimi K3 versus Claude Fable 5 versus GPT 5.6 versus Kimi K2.7 code, Kimi K3 is $3 per input per million tokens. Claude Fable 5 is over three times that at $10 per million tokens. GPT 5.6 is still nearly twice the cost of Kimi K3.
The same with output tokens, right? $50 per output tokens from Claude Fable 5, whereas output from Kimi K3 is way cheaper. I just don't see how they're going to remain competitive with Claude. Like the thing is, if you look at Claude, if you look at OpenAI, they are not profitable, but they need to IPO.
And the problem is that they're getting squeezed by these open source models from China that are absolutely crushing them, number one on benchmarks, number two on actual output, and number three on price, and for me, like I'm a big fan of Claude, love what GPT 5.6 did as well. It's just scary to see how fast these Chinese models are moving as well. Now, also something that's interesting is, as I was talking about before, there was a rapid development and a rapid improvement between last month's Kimi K2.7 and Kimi K3. So for example, we look at the context length here.
You can see this was 262K tokens from Kimi K2.7 code. If we have a look at Kimi K3, that is 1 million context window. It's quadrupled its context window in the space of a month. And it's actually got a slightly bigger context window than Claude Fable 5.
And then if we have a look at the benchmarks, look at the generational gap between K3 and K2.7 code. Like this is not something that's slowing down. This is speeding up and Kimi K3 is the worst it's ever going to be. It's only going to get better from here.
It's only going to improve faster from here as well. So the iterations and the self-improvements just make it better and better. I mean, look at this. Together with refined training and data recipes, it's built on Kimi Delta attention and attention residuals, right?
Which are two architectural updates designed to improve how information flows. And with that whole system, they achieved a 2.5x improvement in overall scaling efficiency compared to K2, which is wild. That's why there's such a big improvement. It's also very good at internal knowledge work.
So it's performing very well. And by the way, some people say, is this better than Opus 4.8? Yes, by a long, long way. On all the benchmarks, I'm seeing it crush Opus 4.8.
I just don't even think Opus 4.8 is anywhere near the same level as K3 at this point. Like if you're still using Opus 4.8, you might as well switch to K3 until they release Opus 5. And I think this actually creates a lot of pressure for Claude to release Opus 5 and have a big step up in terms of overall improvement. And again, this is self-evolving.
So there's actually 15 hours of nonstop iteration and improvement. And K3 just improves itself in its own algorithm doing that. On top of that, it's very, very good at like 3D games. So what this achieves is something called vision in the loop.
So this is really interesting because it's very good at reasoning, coding and vision. It can take its ideas and then self-improve as it goes along. And that's what you're watching inside this video. It can seamlessly iterate between code and live screenshots to keep self-improving the game, for example, which means that if you're creating a game, you could allow it to self-improve and iterate really, really quickly because it can see its own work and then improve it based on that.
Now, bear in mind on the benchmarks, a lot of the benchmarks, just to be 100% transparent here, are still beating KimiK3, so Fable 5 is outperforming it, but you just see it right at the top. It's in the top three for pretty much everything right here. And also when I've tested out myself, you can check out all the tests on GoldieBench. It came out less than 24 hours ago.
We've already created 50 different tests of it. And you can view all those and preview them for yourself. Some of them are still being built, as you can see right here, and they'll just be uploaded later today. But for the actual playable demos, very, very impressive.
I mean, look at this video that he created between Hermes Agent, Reemotion and K3. Like it's a website, but it's a video that's fully animated as he goes along using the combination of Reemotion, a free open source skill, and then using K3 as well. So it's super exciting. I think the competition is good as well.
It creates pressure on Claude and GPT 5.6 to improve massively. You also might be wondering, okay, what are the best ways to use this? So I would plug it into Hermes Agent, make sure that you've got a separate profile for Hermes Agent. We've got it inside our Agent OS, as you can see right here.
You can also plug it into custom workflows. Like for example, we've got this trending tool that basically analyzes all the trending data from Twitter for our industry, and then allows us to automatically in one single click create SEO content and social media content using this, and then we can see the original source that it came up with, and it also runs on a 24 hour schedule using a cron job with Hermes Agent so that it can just find new news every single day and keep me updated. I mean, I can literally just quickly scroll through this and see what are the latest news updates, and then also create marketing content around that topic as well. And the same for Hermes Astros, like this is a custom workflow that you can get inside our Agent OS, where you can analyze your competitors, your keywords, you can watch what they're doing, come up with new angles for your industry, change the keywords that you monitor, change the competitors that you monitor, and then create content inside the Video Agent, Notepad Claim, or SEO content as well.
Plus you can link it to MCPs like you saw with Blender earlier today. Also, we have Kimi Code, and that has a full workspace where everything we create is saved in one place, so we can see all the stuff that we've created. It was actually creating some pretty nice websites, as you can see right here, super nice designs, et cetera. And then also inside the chat, you can use it like a CLI, and you can switch between quality, fast, and no reasoning at all, and this all plugs into a beautiful memory system.
So the great thing about this is we've got an Obsidian memory system, it's a free open source project as well, and you can plug that into your AI agent. So as soon as something like Kimi K3 comes along, or something else comes along, like maybe GLM 5.5 or GLM 6 or something like that drops next, you can just plug your memory system straight in, so you can get the most out of these models and train them up on exactly who you are, what your projects are, your goal, your voice, everything else straight away. And I think those are some of the best ways to use it. Hermes, CLI, and also the memory system.
Also, what I would say over here is the, you could also use it with a combination of Cloud Code. So you could get the agent harness from Cloud Code and then plug Kimi K3 into that. So for example, if we have a look at the APIs plugged in, we've got GLM 5.2 over here, we have GPT 5.6 all over here. We could easily add Kimi K3, and I actually plan to do that later today to test out how it performs with a different agent harness.
So overall, super exciting. It's definitely better than Opus 4.8. It's definitely up there with frontier models like GPT 5.6 and Paypal 5. If you want to get all of our best systems for this sort of stuff and our best trainings, you can get that inside the Airprofit boardroom, link in the comments description or go to the airprofitboardroom.com.
Inside the community, you can ask questions, get help and support. I create a video tutorial answering each of these questions every single day. And then also the whole community helps each other. It's really positive, helpful community of great people learning, growing together.
We have over 200 pages of wins, testimonials, reviews, people learning, growing together. So it's just a great place to see everyone winning and learning. And then also inside the classroom, you get access to all of our best trainings. So we have a complete beginner to expert course on AI over here.
Mostly if you want to get our new daily updates, you can get that with the AgentOS system. We've got a new Conex course as well, a new course on how to create your own AgentOS if you want to start from scratch, or you can just get our system ready to install over here. And then inside the calendar, you can jump on weekly coaching calls, get help and support in real time. Inside the map, you can meet people in your local area who are building with agents like K3, Hermes, and agent operating systems.
And that's all available inside the AI Profit Boardroom. Link in the comments description, or just go to the AIProfitBoard.com. Thanks for watching.
More episodes