The video argues that Grok Build has improved dramatically and is now a surprisingly capable tool, especially for people who already have a Twitter/X subscription since it provides image generation, a CLI coding agent, and no apparent token limits. The creator shows how Grok Build can be used inside an agent setup (e.g., Hermes/OpenClora), including X/Twitter search and web search, and explains its workflow of planning first, splitting work across up to eight parallel sub-agents, self-checking, and iterating before shipping. Examples include building open-world games, simple websites, and an automated “video director” agent that generates fully narrated AI videos with avatars and B-roll. Benchmarks and side-by-side comparisons (Goldie Bench and a Voxel Racer test) suggest Grok Build is solid but generally not as strong as Claude Opus 4.8, and it now lists a 256K context window. The episode ends by promoting the AI Profit Boardroom and its agent operating system, tools, trainings, and coaching calls.
00:00 Grok Build Surprise
01:03 What Grok Build Is
01:58 Games Built Fast
02:20 Speed and Agents
02:56 X Search Advantage
03:31 How the CLI Works
04:15 Best Way to Use It
04:58 Websites and Video Creation
06:08 Benchmark Results
07:12 Voxel Game Comparison
07:55 Context Window Update
08:19 Agent OS and Boardroom Pitch
09:35 Wrap Up
Full transcript
For the first time in my life, Grok build actually seems good, and I've never seen this before. Now, I always wrote off Grok as one of those models that was pretty subpar and tried to do everything with Twitter and everything else. But actually, so far, it's pretty good. I mean, we built this, and this outperformed many of the other stuff that we built using the same prompts on other models, especially the ones that recently came out.
And it's mind-blowing what you can build with this. Plus, if you already have a Twitter subscription, this is great because you can generate images, you can use a CLI, you can code out whatever you want. I haven't hit any token limits. We can build all sorts of crazy stuff with this, and it just comes directly with my Twitter subscription.
So if you're in the same situation, it's pretty good news, right? So I'm going to guide you through what I built with it, how I'm building with it, how I see this stuff going, some of the cool ways you can use this as well, especially agentically. And then what I think, overall, is the future for stuff like this. So let's get into it and stop playing games here.
So if you're not familiar with Grok Build, basically, when you're, you know, if you already got a Twitter subscription, you can just use this and plug this into your system so you don't need to pay for, like, a subscription or APIs or anything like that. And with Grok Build, you can use it inside the CLI, but you can also plug it into your AI agent. So if you go to Hermes, for example, over here, we can use this studio, and I can switch between Grok and Minimax, right? Minimax M3, another Chinese model, pretty good.
Grok Build is inside here, and we can automate images, videos, whatever we want using this section as well. We can even create voice notes, which is pretty cool. And then, not only that, we can actually talk to our agents on a live call, which is pretty amazing as well in itself. So when you're using this, there's actually multiple different ways that you can unlock a powerful CLI without having to do much extra work or pay much extra for it.
Now, is it up there with Grok? I'm going to show you some examples, and we'll compare it side by side. But it is good. Like, I mean, we built this in, like, a couple of hours using Grok, and it was super simple and easy.
Plus, I mean, look at this. This is fully open world. Like, this is a full game that we've built, and I don't code. So how amazing is it that you can build all of this stuff, and you don't need to be a coder, you don't really need to do anything special here, and this game seems to go on forever?
So you can see these two games that we actually built directly with Grok build. And the great thing about this is it scores 70.8% on SWE benches, which is not bad at all. 100 tokens per second when we actually measured it. You can actually have eight sub-agents working in parallel when you're using Grok build, and it's a terminal coding agent as well.
So it's pretty good. Now, Elon Musk was actually asking for feedback on this, and I know why, because it has got dramatically better. Honestly, Grok itself was trash before, if you ever used it. Now, it's a whole lot better, my friends.
And also, what's interesting about this as well is, like, you can use Xsearch. If you plug in Grok into, for example, OpenGlory into Hermes, you can do Twitter search here. You can see some of the searches we've done previously. You can search web, but the great thing is you can search Twitter too, which most LLMs struggle with because they don't have access to that data.
This is, like, the only LLM that actually has access to live data on the most popular news website in the world, a.k.a. Grok, sorry, Twitter, which is pretty amazing. You can also talk to Hermes and OpenClaw, and it's fun to use. It is a lot of fun.
It's got a whole lot better recently. Now, I want to show you an example side-by-side of, like, what we got and how it performs. But basically, you can have, like, eight sub-agents with Grok build working at one time. So you can basically give it an order.
You'll plan it. You'll have, like, sub-agents building in parallel up to eight of them. Then it actually self-checks its work. This is built into the CLI.
And if it's good, it will ship it. If not, it's going to do it again, but sharper, and just keep going around as a positive feedback loop. So Grok builder doesn't just, like, start typing and building. It always plans the job first.
It splits it across sub-agents that work at the same time, checks its own work, then ships, then it loops. And that's why one sentence, which you've seen before, comes back as, like, something super nice, which is pretty awesome in itself. Now, the way that I would use this, because I don't think it's as good as Claude 4.8, is that I'm going to plug it into the agent operating system. And then when I need to build out cool stuff, and I need a lot of tokens, and we need to blast something at our agents, well, we can do that inside the agent operating system.
And that's why I'd recommend for you is, like, you probably wouldn't have this to replace Claude, but you would have it to build stuff, especially if you run out of tokens on the Claude plan. That's a good way to use it. But basically, in one sentence, you can build all sorts of stuff. You have a factory, right?
So in one sentence, you have eight sub-agents working in parallel, and then they build, and then once they've built, you can see that everything lands in your workspace like this, and you can preview it below, which is super useful as well. So it is really, really good. It can even build out websites. Some people say, well, what's the use case of this?
You're not going to be building RPG games all day. But you can see, like, you can build pretty nice, simple websites as well. So it can do all sorts of cool stuff, again, with the images and videos and everything else. Something else that I'll show you, and this kind of blew my mind recently, is that we actually created, like, six-minute-long, fully AI-generated videos that look absolutely awesome.
Let me show you an example of this. So this is an example. We can generate the avatar for the AI video. It can generate B-roll with Grock, and then it has the script over the top, and it's fully narrated with a voice.
Now, this is pretty shocking when you think about it, because it's all put together automatically from one sentence using this whole system here. We call it the Video Director Agent inside our Mission Control. And again, like, when we're creating stuff, it generates all the B-roll for the model with Grock. So that's all automated, and then it uses hyperframes as a skill to plug it all together.
So, you know, real use cases here. Websites, videos. You could do SEO content. You can build out apps.
Like, it is very powerful. It is good stuff. Now, if you're wondering, okay, how does it perform on GoldieBench, which is our benchmark for testing stuff. So again, like, Opus 4.8, I'm always going to choose that.
GLM 5.2, that's won a lot of medals. Quen 3.7, not too bad. KimiK 2.7, I would still say it's not right at the top, right? It's scoring about five on the leaderboard, but has it got a whole lot better?
And is it included for free with many people's subscriptions? Like, if you're on Twitter and you've got a subscription already, is that, you know, can you use it? Yes, absolutely. And you can see, like, the outputs here.
We've compared all the outputs on GoldieBench in terms of how they perform. So this is Grok's, and we can compare these side by side. So we've got Opus 4.8. We have GLM 5.2.
I mean, GLM 5.2 is pretty nice. Opus 4.8 is super nice. Grok is not bad, but it's not the best. I think that's probably where Grok is always going to be from what I've seen.
Like, it's just not bad, but it's not the best. So if you want something, like, super amazing, you're still going to go with Opus 4.8. I mean, look at that. That's goated.
That is beautiful. I could stare at that all day, my friend. Now, if you want to see how it performs versus other stuff, so we can open up this Voxel game and compare it side by side. So this is the one that Grok created, which is pretty nice.
You know, it looks great, as you can see here. It looks pretty cool. It's really fun to play. If we have a look at, for example, the other outputs that we got, this is super boring.
Let's have a look at this one. It's kind of fun, but there's not much going on on screen. Let's have a look at the next one. So on the Voxel Racer test, I would genuinely say, like, I mean, I think this was Opus 4.8.
I would genuinely say that Grok created probably the nicest thing so far, or definitely up there as well. So on a lot of the tests, it's looking really good. And if we have a look here, in terms of how it performs, it's pretty snappy. Pretty small context window, as far as I'm aware, although that might have changed on the latest one.
Let's have a look here. So the new Grok context window is 2K context window, which is not bad at all. And you can see all the demos and all the stuff that's built. So on GoldieBench, you know, if you want to test this stuff, see how it performs, see what we've built with it, you can check it out all here, as you can see.
But yeah, overall, not bad at all. Now, if you want to get Grok along with all of our best systems, so for example, we've actually plugged in Grok, GLM 5.2, Claude, Opus 4.8, Hermes, and all our agents, and a full orchestration dashboard with Paperclip, AI Agent Mastermind, a full pipeline idea to pipeline builds, and Agent Kanban for local engines, then you can get that inside the AI Profit Boardroom, link in the comments description. This is my agent operating system that I've built around Grok and other models like it to get the most out of this stuff. We've even linked, for example, like Notebook LEM, we've got a Kanban board here, full memory galaxy for building with this stuff.
And you can get it all inside the AI Profit Boardroom, link in the comments description, or just go to the AIProfitBoard.com. Now, inside the community, you can ask questions, get help and support whenever you need to. Inside the classroom, you can get access to all my best trainings, including the new daily updates. And we have the Agent OS here that we update literally daily.
So you can see this was last updated today, you get the zip file for it, you get all of our new lessons, and we have new guides and video tutorials every single day that I create to help you personally. Inside the calendar, you can jump on four weekly coaching calls. Inside the map, you can actually meet people in your local area who are building AI agents just like you, which is a great place to connect with people. It's all inside the AI Profit Boardroom, link in the comments description, or just go to the AIProfitBoard.com to get access.
Thanks for watching. Cheers, bye-bye.
More episodes