AI News Today
← All episodes
Episode 86 · July 17, 2026 · 15:31

Kimi K3 VS Fable 5 VS GPT 5.6: Who Wins?

Full transcript

Kimi K 3 vs Claude Fable 5 vs GPT 5.6 wins to date we're gonna be putting them side-by-side to test now so Kimi K 3 is super impressive and I've been blown away by what it can create so far this is actually kind of like a Skyrim open-world style game that we created and it's absolutely massive look at the size of this like how big the full open world game goes then we have Fable 5 which is honestly my favorite before Kimi K 3 came out but let's see what happens after and then we have GPT 5.6 Soul as well so we're comparing all three side-by-side today to see which one creates the best outputs let's kick it off with this Skyrim style game so if we have a look here for example this is really really nice feel smooth I love like the light in the sky the details in the sky I like how open and wide it looks as well it's got kind of that like magical feel to it as well let's have a look if we go inside the village over here you can see it comes up with like the screens as well of where we are and that sort of thing and we've got like all these little houses it just looks really cool looks super nice I mean that's done a great job if we have a look at the Skyrim from Fable 5 I would say the ambience of this just feels a bit nicer also you can see more of the first character when you're going through is a little bit buggy in parts but I would say the Kimi K 3 and Fable 5 have done a great job there when I actually look at the outputs from GPT 5.6 Soul it's not bad but the buttons are totally backwards so if I try to go forwards it's actually the button for going backwards and you can't really go side-by-side so I would say that if I look at all three of those Kimi K 3 and Fable 5 by far did the best outputs now let's have a look at a similar sort of game called Dragon Realm that we test with all of our agents by the way if you're wondering okay how relentlessly did you test this model so I mean it dropped less than 24 hours ago and we've actually tested it out across 50 different tasks so far so we have tested it relentlessly it does create some awesome stuff and I'll come on to it in a second so let's have a look at this this is the output from Kimi K 3 which looks absolutely awesome doesn't it like the snow falling the graphics look at the sky it's wild moving around it's pretty easy easy to control and we've got this dragon over here which I do know I've never seen any of my AIs create that style of dragon on a normal test like this so that's the first time I've seen an AI model complete the test like that which is pretty impressive in itself let's have a look at Fable 5 now this does feel a little bit more basic than Kimi K 3 in terms of the details and everything else genuinely I would say K 3 did a better job of Fable 5 here and then if we have a look for example we've got GPT 5.6 GPT 5.6 I would say probably did slightly better than Fable 5 especially with adding the enemies and that sort of thing so I'm gonna go with K 3 for winning on detail and ambience GPT 5.6 did a awesome job in just making the gameplay a bit more interesting and then Fable 5 I would say we pretty much struggled on that test to be honest now we have a racing game we can just compare these side by side here so if we open this all up we've got this one from Kimi K 3 looks awesome feels nice I just love the colors and the vibe and the gameplay is smooth I'm noticing that on every single game it's so hard to describe but it feels smooth it feels nice to use and I think that would be amazing for creative websites as well now if we have a look at Fable 5 this has not done a bad job but it just feels a bit more basic it's a bit more like a block sort of thing so the graphics over here are nicer the details over here are nicer from K 3 and then if we have a look we've got GPT 5.6 which I would say is done better than Fable 5 it's not as interesting to play but the graphics themselves look nicer than Kimi K 3 but again there's nothing to dodge the gameplay is not really there it's not really that exciting and it slows down a little bit as well so I'm going to go with K 3 winning on that one too let's have a look we've got the next one which is Neon City we've got a lot of Neon tests coming up so we have a look at this I mean this is insane this is K 3 over here looks awesome feels fun to drive pretty intense but we love it love the detail of all the city blocks and everything else the colors everything else if we have a look at Fable 5 it just feels more basic it almost feels like a generation behind on that one and then if you look at GPT 5.6 it's okay it's better than Fable 5 but it's not quite as good as K 3 so K 3 is really really powerful at this point and it can create some amazing stuff now if we look at the crit game this kind of feels a bit more like a maze there's some nice lighting I like the graphics I like the details it feels nice to use it if we have a look at the version from Fable 5 this is a bit more linear but it's a it's a bit more interesting in terms of gameplay so when we're doing the tests here you can see it's also quite buggy breaking a lot let's have a look at the next one this is similar sort of vibe but from GPT 5.6 which I would say again is beating Fable 5 there so I think in terms of graphics K 3 in terms of gameplay GPT 5.6 that's quite interesting there you know the AI's are good at different things aren't they one can be great at graphics one can be great at gameplay you take your pick and decide which one you prefer I think on this one K 3 actually failed so we might have to regenerate that later you can see this just doesn't work if we have a look at GPT 5.6 and Fable 5 over here very basic output from Fable 5 but it's it's playable it can do the job the one from GPT 5.6 looks 10 times better than anything else here like look how cool this is the graphics the colors the gameplay everything is much better now this is a black hole simulation I would say definitely K 3 has created the most visual interest in one here these two just feel a little bit not quite right like something is broken inside there or it doesn't understand the problem properly or something like that whereas if you look at K 3 like super visual nice we can move it around we can preview it side by side looks way nicer than the black hole simulation from GPT 5.6 and Fable 5 so Kimi K 3 you know I think we're really at the point now where open source models are right at the frontier you know and they're only going to improve faster I would expect to see a new model from Kimi and GLM every single month for this rate the rate that they've improved the progress over the last few months they're evolving way faster whereas you look at something like Claude Fable 5 you look at GPT 5.6 and there's a lot of back and forth number one in terms of getting the models out there sometimes they get taken down then they come back out and also bear in mind that GPT 5.6 and Claude Fable 5 you have to be very careful with tokens so for example when I was running the Goldie Bench experiments we had to use the API on a lot of the tests because it just ran out of tokens on the subscription with chat GPT and Claude that has not happened with Kimi K3 and we're just on like the $39 a month plan we've had to regenerate some of the tests and it still was just just steaming a lot so the thing is not just that the quality is right up there with both of these models but also that Kimi K3 is open source it's from China and additionally it's way way cheaper on the token plan so if you get the coding plan you can actually plug it into your Hermes agent which you can do with chat GPT but you can't do with Claude Fable 5 unless you use the API which gets expensive. Now we have this fluid in a box test looks pretty crazy from Kimi K3 Fable 5 just created something really basic there GPT 5.6 did something interesting but for sure like Kimi K3 is just crashing on these tests now let's talk about benchmarks by the way for me personally I don't really pay attention to benchmarks so much I like to test this stuff out myself particularly when it's not a company that I'm massively familiar with so I like to test out myself and you've seen on the benchmarks actually in reality K3 is outperforming a lot of these different models but if we have a look over here so if we look at this we've got for example Kimi K3 and that is being outperformed by Fable 5 and GPT 5.6 so on DeepSWE and again like I just pay attention to my own benchmarks that's why we create GoldieBench because I just don't think that these benchmarks are realistic sometimes they're too optimistic sometimes they're not optimistic at all either way test yourself out for yourself don't even listen to me you know just make up your own mind if we have a look over here as well terminal bench 2.1 so we've got Kimi K3 versus GPT 5.6 so Fable 5 is way down here at 84.6 Kimi K3 scores 88.3 GPT 5.6 scores 88.8 these are according to the benchmarks that have been released by Leo over here thanks to Leo I've not seen the official benchmarks let's see if we can find them so they put the API documentation here as you can see but there's not much information on the benchmarks I can't find them on the website either however we can have a look and compare them on Maruta so let's compare Fable 5 and also GPT 5.6 Sol by the way Sol is the model that you want to compare against all of these so if we have a look here they've all got the same context window Moonshot AI is a Chinese lab which created Kimi, Anthropic is the founder of Claude, OpenAI is founder of GPT 5.6 and they've all got the same context length they're all reasoning models they all have very similar input and output modalities as well the price is a big big difference here right now actually what's surprising here I guess it's not that surprising because models have just improved so much but if you look at the comparisons Kimi K2.7 was actually a lot cheaper so let's add that to the list as well so yeah I mean look at the difference in price here like Kimi K2.7 which is the previous generation of Kimi K3 was way way cheaper for example input tokens for Kimi K3 is three dollars per million tokens whereas for example Kimi K2.7 is 0.72 cents however the step up in quality is probably about two to three times better like it's way more powerful it would also be very interesting to use Kimi K3 with GPT 5.6 as a mixture of experts models so for example we've got mixture of agents with Fable 5 and you can compare them side by side but I would say that would be an amazing way to get better outputs like fusing the models together and just seeing how the agents perform when they work together as a team in terms of latency Kimi K2.7 code is a lot faster than K3 and also bear in mind Fable 5 GPT 5.6 not open source Kimi K3 is open source as well so let's have a look at the benchmarks now so Kimi K2.7 is a huge gap in terms of the improvement here so you've got Kimi K3 holding its own with all the frontier models Kimi K2.7 was nowhere near and what actually surprised me here is like how quickly Kimi K2.7 improved it was only like last month that it came out now it would also be interesting to see what Kimi K3.7 is being used inside so for example where are people using it the most I can imagine it's gonna be Hermes as number one ah Claude Code so if we have a look at this Hermes agent is actually third on the list Claude Code is the number one place and the number one app that people are using Kimi K3 in so that will be interesting experiments as well as like plugging Kimi K3.7 into Claude Code and just seeing how it performs inside Claude's agent harness which you can easily do and oh my pie this is not an app that I've tested out much but I'll be excited to test out with Kimi K3.7 as well might be something in the future that we do and I think one of the best ways to get the most out of all this stuff so we've got Claude Codex and Kimi Code all plugged into our agent operating system and that means we can combine the both we can have a group chat where all of our agents including Kimi K3 speak over here we have for example paperclip where we could orchestrate all three agents together as a team inside a company and then for example we have the memory system so as soon as Kimi K3 came out we were not starting from scratch because it has all of that memory and all these different pieces of context about me inside a beautiful memory galaxy using Obsidian so I think that's one of the best ways to use it we also have a way of basically switching between GPT 5.6 Sol and GLM 5.2 inside Claude Code and I'm probably going to add Kimi K3 inside there as well so those are some interesting ways you can use this as soon as the rest of the results from Kimi K3 come out on Goldiebench as well we'll show you those they're just being built in the background that's why they're not done bear in mind this is only just dropped 24 hours ago something else that you can do with Kimi K3 is you can build out these beautiful videos so this is a video that we actually created using Kimi K3 with ReMotion as a skill and it created like a really nice video nice camera angles nice animations etc fully automated in the background and it's actually it actually moves as a background when we move our mouse which I've never seen with a video before so it's kind of like overlaid with a website and a video at the same time so which one would I pick overall I would say for 3D games honestly like K3 outperformed all of those different games I also think that with the coding plan it's a lot cheaper to use K3 for me personally if I'm building out something big like the Agent OS where I can't make any mistakes really then I'm going to stick with Claude because we have our systems built into that and also it's fully trained on all the skills that we need but I think K3 is a fantastic cheaper alternative to Fable 5 and GPT 5.6 so if you're wondering about the best way to get the most out of all of them I would make sure they have an agent operating system with all of these agents side by side so if you want to get our agent operating system you can get that inside our AR Profitable Community this is a place where you can learn grow and scale with AR to make sure we have all of our best trainings inside here you get our video tutorial for the Agent OS system you can see that it was just updated today with K3 and then you can also get the zip file to install it and we had new daily tutorials based on what's just dropped additionally inside the community you can ask questions get help and support in real time and I personally answer all the questions inside the community with a video tutorial and then inside the calendar you can jump over with coach goals get help and support in real time share your screen meet other members inside the map you can actually meet people in your local area who are building out with AR agents just like you and that's all available inside the AR Profitable Boardroom link in the comments description or just go to the AR Profitable Boardroom.com thanks for watching

More episodes

Browse all episodes →