Kimi K 2.7 Code has a new high-speed rollout with three “gears” for the same coding model: Quality (thinking on), Fast (thinking on but served faster), and No Think (thinking steps off for maximum speed but sometimes lower quality). The video compares side-by-side outputs and response times on three prompts—building a 3D solar system, generating a playable game, and creating a spiral galaxy—showing mixed results: some tasks are faster with No Think (e.g., 98s vs 53s vs 39s), while others can be slower, depending on what gets built. The script explains how thinking-on uses a hidden scratchpad while thinking-off writes immediately, and notes Moonshot’s quoted ~180–260 tokens/sec serving. It also demonstrates toggling modes in KimiCode within an Agent OS setup and promotes the AI Profit Boardroom for tutorials, tools, and coaching.
00:00 Kimi K 2.7 Speed Update
00:25 Three Gears Explained
01:22 Solar System Speed Test
02:26 Playable Game Comparison
03:16 Spiral Galaxy Surprise
03:58 Switching Modes Demo
05:12 Thinking On vs Off
06:27 Which Mode To Use
07:00 Quality Mode Showcase
07:46 Agent OS And Boardroom
08:57 Wrap Up And Thanks
Full transcript
So Kimi K 2.7 code is just released high speed. This is a brand new update from Kimi and basically this is rolling out right now. It's up to six times faster. I actually tested it and I tested it across the non-thicking and also the other version of Kimi K 2.7 code just to see how it compares and I'll show you what we learned today and what we've got so far.
So this is basically how it works and there's basically like three different versions of Kimi K 2.7 code now right. So there's three different gears. That's the way that I see it. It's one model but three different ways to run it.
So before this you know Kimi K 2.7 is one coding model. What changes now is which lane you use. So you've got the quality mode model which is Kimi for coding and that's like the full model with thinking mode on. It works like the problem out first then rise it and that's a careful one.
That's not the new update. I mean it came out two days ago but it's not the new new update. Then you have fast mode which is the high speed lane which is a faster way to serve the same model but with thinking mode on and then you actually have thinking mode off which is super high speed right but the only problem is like the think out loud sort of steps to switch it off. So it answers straight and answers quickly but it's it's not as good but let's see what the outputs are.
So we've actually got some tests here in terms of side by side what we got using the same sort of prompt. So this was the prompt which was build a 3D solar system the sun plus all eight planets orbiting at different speeds with small labels and you can see the outputs here. So this is like the normal sort of quality model and then this is the fast mode with thinking mode on and then this is no think. Now we actually compared the speeds of each response as well so we measured those.
We got 71 seconds to build this with quality and thinking mode then fast was 65 seconds and then no think fast was 47 seconds. So it's almost two times faster than the quality mode and again like it depends what you build but it can be up to six times faster. Having said that if you look at this they kind of look similar but the way that they're built looks slightly different. There's not a huge difference between each of them maybe someone who's like a bit more science minded will be able to judge on that to be honest but let's have a look on the next one.
So the next one we gave you a complete playable game again tested no think versus fast versus quality so if we actually look at the speed see the quality responded way faster to build out which was weird so this took 125 seconds with no think which is kind of weird so sometimes it's like actually slower using the no think which I don't understand how that's happened but it did we you know we measured this directly so these are the games that we built as you can see right here and they're all pretty much the same but that was interesting too. Now all three are genuinely playable but fast and no think both wrapped the play field in a 3d framed border which adds more depth so you can see the 3d border over here I think that's why it was a bit slower to respond which is quite interesting so it's not always faster basically that's the whole point here and then it and it really does depend on what you build also created a spiral galaxy here I actually think like if I look at these fast mode with no think created like the nicest version right like that looks the nicest out of all these if I had to pick one I'd go with that one so sometimes no think as well it can reply faster so 39 seconds versus 98 seconds for the normal mode but it creates something better as well it's really really mixed depends what you're building and that sort of thing but I've not seen like a 6x speed increase is unless you're using like a super basic question if you look at this night set 98 seconds versus 53 seconds versus 39 seconds here so those were three tests that we ran with this now you might wonder okay like what actually happens under the hood you know what's going on here how does it work so the way that we've set this up you can use this however you want but the way that I've built with this is we've got Kimmy code here and then we can switch between quality fast and no think so these are basically toggles so we can click between them and then we can use the agent here so if we say okay you're here inside the chat Kimmy is thinking and this is like normal mode if we use fast mode it should be a lot faster to reply and then if we use no think that should be even faster so that's how you can switch and also it applies like slightly differently on no think versus fast and thinking mode so those are like the same two replies with thinking and thinking fast mode so here's how it works you got like one prompt it's going to think first with the quality mode then it's got standard lane then it's going to build with fast mode it thinks high speed lane and then it builds and these two think as you can see and then with no think it doesn't think at all reasons just in line when it's coming up with the answer builds at high speed and then builds so you've got like different switches depending on what you want so quality and fast both think first they just run on different lanes now you might think okay what's the difference here so thinking on versus thinking off so we're thinking on on an AR model and this is useful for like understanding how any AR model works but basically we're thinking on the model works the problem out on a hidden scratch pad first plans the layout the physics and maths then writes the code we're thinking off there's no scratch pad so it just starts writing straight away now off does not mean it skips to thinking the reasoning just moves into the answer itself and that's what usually costs time now with high speed it's a faster way to serve the exact same model so moonshot quotes around 180 tokens second up to 260 tokens on short jobs which means like you just get the response way faster so it's the same brain the same training just push for a faster pipe you can see the differences right here and you can see the median build time across all of our builds so the three tests i show you you know 47 seconds versus 65 versus 85 so on the tests i've seen it's about half the time to get a response from no think but also with fast mode as well it swings like really wildly so you can see the difference here this is fast mode and it can take anywhere between like 47 seconds and 188 seconds to reply and the way that i would look at this is like there's no best gear so to speak so there's the right gear for the job you know which is the whole point of the framework so for example if you are using really fast mode this just for like quick jobs you know throw away tests simple utilities where the the polish doesn't matter then you've got fast mode which is great for like iterating i would say and like just improving the projects you already built and then you've got for example the builds that really matter or like just building something from scratch and that's where i'd use the quality mode and like it can build pretty cool stuff so if we have a look for example this was some stuff that we built with kimi k 2.7 using the quality mode so you can like create videos as you can see here create avatar videos it can for example create games all sorts of cool stuff here even built out like a full world game here so if we have a look at this this was like a full open world game that we built using kimi k 2.7 and then we also created this one too which is pretty fun to play so this is another example of like what we've built with kimi k 2.7 so so far it's pretty good uh the fast mode again i would just use it for like super quick answers or just improving what you already got i wouldn't use it as a default some people will say well faster code means worse code but you've already seen like some of the projects we built with fast mode are actually better so that's basically it now if you want the whole system that we've built out for using kimi k 2.7 which you can see here so we've got kimi code built into this agent operating system it's using the cli it's your quality fast and no thing that you can switch between you've got all your agents inside one mission control you can use hermes and claude and whatever you want to build with over here you even got free claude code built in and grok build and fusion which is a way to achieve like fable fable 5 level intelligence without having access to fable 5 so you can get that all inside the agent os system which we've got inside the ai profitable room so if you check it out link in the comments description or go to the air profitable.com and then you go to the classroom here new daily updates you'll find the agent os system with the last update here we have a full tutorial and roadmap on how to use kimi k 2.7 glm 5.2 is a good option too and we basically show you everything that you need to learn about ai automation right now you can ask questions inside here you can go inside the classroom get all of our best trainings you can jump on with coaching calls you can meet people in your local area as well who are using ai agents like you and this is all inside the air profitable room link in the comments description or go to the air profitable.com thanks for watching
More episodes