AI News Today
← All episodes
Episode 59 · July 1, 2026 · 07:39

China’s NEW Meituan LongCat 2.0 Tested!

LongCat 2.0 (Open Source) Tested: Benchmarks, Games, and GLM 5.2 Comparison

The episode covers the official release of LongCat 2.0, an open-source Chinese agentic model revealed as the model behind the AoAlpha free API, with features like Sparse Attention, Zero Compute Experts, and MIPD. The host reviews benchmark claims (including Terminal Bench 2.1 and SWE-Bench Pro comparisons versus GPT-5.5 and Opus 4.8) and shares hands-on tests building game demos such as Dragon Realm, a Skyrim-style open world, and VoxelCraft, noting mixed results and frequent bugs. Access issues are mentioned, including difficulty using the API without a Chinese setup, so the model is tested via the website chat. A key point is that LongCat was trained on China’s Meituan chips without NVIDIA. Overall, GLM 5.2 is judged stronger in side-by-side game benchmarks, and the host promotes the AI Profit Boardroom and Agent OS setup.

00:00 LongCat 2.0 Launch
00:36 Benchmarks and API Hurdles
01:38 Game Demos Dragon Realm
02:23 Goldy Bench Verdict
02:43 Trained Without NVIDIA
03:32 How to Use It
03:51 Eval Results vs GPT
04:17 GLM 5.2 Showdown
06:13 Final Take and Recommendation
06:35 Agent OS and Boardroom Plug
07:37 Wrap Up

Full transcript

Today, we have a brand new update from a Chinese model that's open source. It is called LongCat 2.0 is here. And this was actually the full model behind our alpha. So if you're familiar with our alpha, which was a free API, you could actually use it with Hermes.

You could plug it into code code before, and it was not bad at all. It's a genetic model. And this was actually revealed as LongCat 2.0. So this is now officially been released and you can get access to it and you can see the full details right here.

So it's got LongCat sparse attention, zero compute experts, MOPD stacks up not badly on the benchmarks here. If you're wondering how to compare against everything, so terminal bench 2.1, it holds its own with these other models, as you can see right here. This compared against Opus 4.6, 4.7 and 4.8. Now, obviously Opus 4.8 is crushing it by a long way, as you can see here.

Um, the other thing that I noticed is if you're using this directly on the website. So if you go to LongCat and then you go to the API section, it looks like you can't use the API unless you have some sort of Chinese setup. So for example, here, if we try and get a pack, it's broken anyway, look at that. Um, but if, if you want to use it, I've been testing out.

So I'll show you what we've done so far with it and how it works and how it performs here, so let's have a look at some of the stuff we've built with this. How perform some benchmarks, some of it was good. Some of it was not so good. So this is a game we created called Dragon Realm.

It's, it's not bad. I mean, the graphics are not bad. This is nowhere near the same standard as GLM 5.2 or Opus 4.8 or anything like that. But I don't think it's designed to be, I think it's just designed to be a cool model that you can test out.

This one was not bad at all as well. So it's kind of like a Skyrim style open world game. As you can see, it runs pretty smoothly. The graphics are not great.

And then here's another one. This is probably the best output that I saw so far. And we can compare it on GoldieBench against a bunch of other examples in a second, um, but again, it's still pretty basic and buggy, right? Like, like what is going on here?

It just goes completely black. And then we have another example right here. So overall on my tests, I wouldn't say it's that impressive. We've ran it through GoldieBench.

I'll show you how it compares versus our models in a second. It is a 1.6 open source, 1.6 trillion parameter model. And the other thing to note here, this is probably the biggest update about it. As you can see from this tweet by Robin, is that number one, it's open source, but number two was actually built on a different chip.

So this was built on China's Meichuan. And this was trained without a single NVIDIA chip. Now, if you're wondering who are Meichuan, they're basically like China's version of DoorDash. This is pretty crazy.

I mean, we saw this with Xiaomi as well earlier this year, where like, you know, everyone's getting involved in AI, everyone's bringing out their own models. Companies that we don't expect to bring out their own models come out. But yeah, it's pretty cool. I mean, like fair play to them for bringing out a model like this.

Would I say it's the best model I've ever used? Probably not, but it's fun to play with, fun to test out, et cetera. The way that I actually used it, because I couldn't get access to the API and see like it's, it's not even available to top up yet, is you can actually go to the chat section and just start using it there. They've also got a full breakdown of how it works step by step, as you can see on the website.

Also on the evaluations, this is pretty interesting. GPT 5.5 is only slightly above on Terminal Bench versus Longcat. And then on SWE Bench Pro, Longcat is actually outperforming GPT 5.5. Now again, I've shown you my tests.

Do I personally think that it's better than GPT 5.5 on my own benchmarks? Probably not. And Opus 4.8 is beating them all, but you would expect that anyway. Now, I think probably the most relevant side-by-side comparison is comparing it against GLM 5.2, because that is another open source project that just came out of China recently, and it's very cheap on the coding plan, and also it's designed to be agentic as well.

So for example, we have a look at this Crypt game. This is the one created from GLM 5.2. It's pretty dark and it is a little bit buggy, but it works and it's actually, you know, interesting and quite useful. If we compare that versus the option from Longcat here, you can see that this is very limited.

It has nothing going on in the game and it's super buggy, right? You can just walk through the wall. So on that particular benchmark, GLM 5.2 is winning. Let's have a look at the next one.

So this is the Dragon Realm game, as we showed earlier. Now this is the one from GLM 5.2. Pretty nice. Open world, has lots of interesting stuff, quite playable, et cetera.

Right? And the, even like the graphics are not bad at all. Now, if we compare that versus this option from Longcat here, you can see that it's just not quite the same quality or caliber, it kind of feels like this is an older generation of it. Let's check the next version now.

So we've got our own version of Skyrim over here. This is the GLM 5.2 option. Again, looks pretty nice. Very open world, feels expansive and big.

If we compare that versus Longcat here. This is actually a lot smoother, but there's just not much going on, right? This is not as interesting. And then we created Voxel Craft as well.

So this is the option from GLM 5.2, which is actually pretty impressive. Like it's, it works perfectly. It's fun to play. It's easy to set up, et cetera.

If we compare that versus the version from Longcat, you can see it's just nowhere near the same level. So if you have a choice between them, I would still go with GLM 5.2. I think that it's interesting to see what's coming out of China, but I think GLM 5.2 by far is the strongest model out of everything I've tested when it comes to open source stuff. So that's basically it.

Will I be switching to Longcat anytime soon? No, but I wanted to test it, show you what it looks like, et cetera. One thing we've actually done with GLM 5.2 recently is plug it into Cloud Code inside the agent operating system. And this is our mission control dashboard where we can have all of our agents working and plugged in together.

Now, if you want to get that set up, it's all available inside the AI Profit Boardroom, link in the comments description, or go to the AIprofitboardroom.com. Inside the community, you can ask questions, get help and support in real time. I answer every single question inside the community, personally with a video tutorial, plus you get the whole support and help of the whole community. And there's always someone online 24 seven to help you.

Inside the classroom, you can get access to all of our best trainings and lessons. Inside the new daily update section, you can see we've got our full HNOS system, the last update date, so we update this daily, the video tutorial, and a zip file on how to use it. And then we add new tutorials based on what actually comes out and is useful. Inside the calendar, you can drop off weekly coaching calls, get help and support in real time.

Inside the map, you can meet people in your local area who are building with AI agents like you've seen today. And that's all available inside the AI Profit Boardroom. Link in the comments description or go to the AIprofitboardroom.com. Thanks for watching.

More episodes

Browse all episodes →