AI News Today
← All episodes
Episode 36 · June 16, 2026 · 09:03

NEW GLM 5.2 BEATS Claude?

GLM 52 vs Qwen 37 Max vs Claude Opus 48: Real-World Tests vs Benchmarks (No Second Chances)

The episode compares GLM 52 (ZAI), Qwen 37 Max (Alibaba), and Claude Opus 48 (Anthropic) head-to-head on five one-shot tasks, arguing that benchmark rankings didn’t match real usability. In coding-focused tests like a voxel runner game, a liquid-in-a-bowl animation, a business landing page, and an arcade game, GLM 52 produced the most fun, polished, and feature-rich results, while Claude’s outputs were often basic and Qwen’s were sometimes buggy or incomplete; Claude clearly won the solar-system orbit map task. The script also notes Qwen’s strong reported benchmarks and faster replies, GLM’s slower responses in agents but strong CLI coding, and highlights limitations integrating Claude into agent workflows compared to Qwen/GLM in Hermes and the creator’s agent operating system.

00:00 Head To Head Setup
01:27 Coding Tests Results
04:09 Arcade Game Showdown
04:50 Benchmarks Versus Reality
06:01 Agents Workflow Tradeoffs
07:59 Final Recommendations

Full transcript

GLM 5.2 versus Quen 3.7 versus Claude Opus 4.8. I put all three head-to-head, same five tasks, one shot each, no second chances and the results flipped everything I thought I knew about picking an AI. So here's the part that gets me. The model that wins on paper the best one with the best scores came in last when I actually used it and the one that looked the best that actually shipped with no official scores at all.

So if you've been sitting there trying to figure out which AI to actually use for your business and you keep seeing these big benchmark charts that say you know something different every single time, well we're going to look at what performs the best out of Quen, GLM 5.2 and Claude Opus 4.8. Now these are three of the top AR models right now. GLM 5.2 is from a Chinese company called Jiabu ZAI. There's a coding plan then you got Quen 3.7 Max from Alibaba and Claude Opus 4.8 from Anthropic.

Now I ran all three through the same jobs and we'll start this off. So the first job that we gave it here was a sort of voxel runner game. So we've got GLM 5.2 over here, Quen 3.7 and Opus 4.8. So this is the one from Quen 3.2, sorry from GLM 5.2 as you can see here.

It's a lot of fun, pretty interesting game, pretty cool and a lot of fun to play. If we look at this one from Quen 3.7, this is Quen 3.7 Max by the way and you can see here that it's quite boring to play with right. I mean look at that, it's kind of buggy but it does the job, it does the job. Then we have Claude Opus 4.8 and look how basic that is.

That is not so much fun at all. So on the first coding test here we can see clearly the GLM 5.2 is winning and Quen 3.7 comes in second. Next up we have the inner system orbit map and actually if you look at all three of these undeniably Opus 4.8 wins this right. It's doing the best job here and it's done something amazing as you can see and we can change the speed, we can change the rotations etc.

This one is okay, this one is pretty bad. Now you can like zoom in and it looks better and that sort of thing but on the outside it doesn't look that great. I mean it's kind of cool to play with but I think out of all of these you know Claude Opus 4.8 absolutely nailed it. Now we've got the liquid in a bowl test.

So we have Quen 3.7 over here, GLM 5.2 and Opus 4.8. Now if you look at these, I mean this is kind of a boring test but you can see here that the animation from GLM 5.2 is really nice. Like we can change this, we can change the theme, it's pretty cool to play with. We have a look at for example the one from Quen 3.7 max, not quite as fun, it just kind of you know fades out very quickly.

Then if we have a look at the one from Opus 4.8, look how boring that is compared to what GLM 5.2 created. This is way better, way more fun, way more interesting and that's what we want really. So on test 3 and 1, GLM 5.2 won and then on test number 2 which is the Galaxy Orbit you can see that Opus 4.8 won. Then we've got the landing page test.

So this is useful if you're checking for example you know if it's actually creating something useful for business, so creating a website. Now let's have a look at this. This is Opus 4.8, super basic, super boring, not much to it at all, not that interesting. If we have a look at Quen 3.7, it's okay.

I mean this is kind of weird because there's nothing here right, it's kind of just like a empty canvas, but the rest of it was okay. Then if we have a look at GLM 5.2, if we scroll down it's got some nice animations, there's a lot more to the page, nice nicely but cleanly as well set up and I like even like the animations on the page that look super nice and you see how it's actually filled in the canvas whereas for example Quen 3.7 didn't do anything and the one from Opus 4.8 is super boring. So that's the difference, this one, the landing page test as well. Then we have the arcade game.

So this is pretty cool, pretty fun from Quen 3.7 but the only issue is you see how the ball doesn't actually bounce off the walls, like it just disappears completely. If we have a look at Opus 4.8, this has built something better that is more playable and more useful and then if we have a look at GLM 5.2 here, look how cool this is. This is way more fun and interesting and so GLM 5.2 won on pretty much all of the tests apart from one, which is mind-blowing in itself and so in terms of the actual tests I've run here, I would say the GLM 5.2, which is a new model from ZAI, is actually beating Opus 4.8 and definitely beating Quen 3.7 Max. Now if we actually have a look at the benchmarks here, Quen 3.7 Max is the strongest of all three, you know, Alibaba reports 80.4% on SW Bench Verified.

We don't have the benchmarks, we only have GLM 5.1 benchmarks, so it's not really that useful but if we're comparing side-by-side Quen versus Claude, well it's actually beating Opus 4.7 on agentic coding. However, one thing to note here is that Opus 4.8 was so new when Quen 3.7 Max came out that they didn't include Opus 4.8 on their benchmarks as well, so pretty interesting. For me personally, do I really pay much attention to benchmarks? No, I just test stuff in reality because, you know, just how can you believe in the tests run by the company that owns it and also do benchmarks from my experience always translate into what you see in reality?

No. So there's a big difference here but the main thing I would say is like, you know, China's GLM 5.2 so far has been really really good and if we have a look inside the agent operating system too, we can check the workspace and see what we've built out here and we created some awesome stuff as you can see here too. I mean, this is like a full open world game, this is another one kind of like a Skyrim style RPG, this was a fun one too as you can see here right and this was all created using GLM 5.2. So it's a pretty powerful model and the other thing I would say here is like you can actually plug Quen 3.7 Max directly into Hermes agent and you can do that with GLM 5.2 as well on the coding plan but you can't actually do that with Claude.

So Claude doesn't allow your subscription to plug into your agents whereas for example we can create separate profiles for GLM 5.2, for Quen 3.7 Max and for using Kimi K 2.7 which is another model. So when you're on these coding plans you can easily swap them in and out of your agents which I think is super useful in itself as well. The other thing I would say here is like when I've been using Hermes with GLM 5.2 it seems super slow to reply Quen 3.7 Max does seem a lot faster and also if you look at the quality responses here, so we asked GLM 5.2 to take a look at our Obsidian memory, it wasn't as useful as when we actually got the answer back from Quen 3.7. Quen 3.7 actually reviewed all of our notes, looked at what we were ranking for and then gave us some great SEO keyword research.

When you check GLM 5.2 it's super brief and not as useful so it's pretty interesting to see like okay if I'm using AI agents like Hermes I might not use GLM 5.2 but if I'm coding directly inside the CLI for sure I'm going to use GLM 5.2 or Kimi K 2.7. These are great models have created some awesome stuff. So it depends how you're going to use it as well and another thing to note here is we actually got a team of video agents to work together and again you can't really do this with Claude unless you go to Claude directly but with Hermes you can easily get them to work together and then create some awesome stuff. So if we go to the Kanban board here we got a team of agents to work on some videos and fully create it from scratch and then if we open this up this is the finished video as you can see and this is fully AI generated.

I just gave it the prompt and then we had the judge tell them to iterate and keep going until it was finally done but yeah again pretty powerful model for having your AI agents work together. So thanks for watching, if you want to get my full agent operating system with everything plugged in and we have multiple models so we have all of these working together all three coders in one place and I would recommend that for you because then you get the best out of all of these models depending on their strengths. So if we're actually coding on the back end like I tend to find the Claude is better for just pulling them all together and orchestrating them. If I'm trying to create something really awesome and cool in terms of like an actual coding project like you saw with the games that we created then I'll probably go to GLM 5.2 but the main thing is you want all of them in one place and so if you want one dashboard we've actually got the agent operating system with one dashboard, one workspace, you can preview everything live, you can see it built, you can see them all working together with the memory system, the memory galaxy, we've got obsidian and that's all inside the AI Profitable Boardroom.

Link in the comments description or go to the AI Profitable Boardroom.com. Inside the community you can ask questions, inside the classroom you get access to all my best lessons, inside the calendar you can drop on weekly coaching calls. Hope to see you inside, cheers for watching, bye bye.

More episodes

Browse all episodes →