AI News Today
← All episodes
Episode 7 · May 29, 2026 · 15:12

China’s Qwen 3.7 Max DESTROYS Claude?

Qwen 3.7 Max: Alibaba’s 35-Hour Autonomous Agent Demo + Claude-Beating Benchmarks (With Caveats)

Alibaba’s new flagship model, Qwen 3.7 Max, was unveiled around the Alibaba Cloud Summit in Hangzhou (May 20, 2026) and is positioned as a closed, proprietary frontier model aimed at enterprise, narrowing the gap with Claude Opus 4.7 while costing less per token. The script highlights strong agentic benchmarks (e.g., Terminal Bench 2.0, SWE-Bench Pro, MC Atlas, GPQA Diamond) and broad compatibility with agent frameworks and APIs (OpenAI and Anthropic specs), plus availability across multiple platforms. It also stresses caveats: the model is unusually verbose, which can raise real costs, and it has a low hallucination rate partly due to a much lower attempt rate. A headline 35-hour autonomous optimization demo (vendor-stated, not independently verified) reportedly achieved a 10× speedup on Alibaba’s Shenwu M890 chip kernel.

00:00 Qwen Shocks The Frontier
01:34 What Qwen 3.7 Max Is
02:17 Agent Framework Compatibility
02:55 Benchmark Wins Explained
04:12 Pricing And Token Trap
05:29 Hallucinations Versus Refusals
06:27 Inside The 35 Hour Demo
08:19 Hermes Agent Integration
09:50 Should You Switch Now
11:39 How To Test And Deploy
12:36 Stop Waiting Start Building
14:31 Final Takeaways And Caveats

Full transcript

Alibaba's Kuen 3.7 Max just ran for 35 hours straight. Not just that, but it's also destroying Claude in many benchmarks. This is China's most popular model right now across the world in terms of being an agentic tool. And this was running for 35 hours straight with no human, no breaks, 1,158 tool calls.

And at the end of it, it had an optimized complex AI kernel to run 10 times faster than the original. This is not a marketing slide. That is the demo they showed at the Alibaba Cloud Summit in Hangzhou on May the 20th, 2026. And it should have way more people talking about it because here's what actually happened.

Alibaba just released a frontier-level AI model, Kuen 3.7 Max. It scores within 0.7 points of Claude Opus 4.7 on the Artificial Intelligence Index, the most credible independent AI benchmark out there right now, and it costs half the price. Not just that, but it's actually beating many of the benchmarks for Claude Opus 4.6 as well. So if you've been spending money on Claude Opus 4.7 or GPT 5.5 for your business, you need to know about this because the AI landscape in 2026 is not just anthropic and open AI anymore.

China just showed up at the frontier. So let me talk you through exactly what this model is, what it can do, where it beats the competition, and crucially, where it doesn't because there are real caveats to this, of course, as well. And I'd rather you know them up front than find out the halfway. So Kuen 3.7 Max, what is it?

It's Alibaba's new flagship AI model. And this is a big shift for Alibaba because they've always been known for releasing open-source models most anyone can download and run themselves, Kuen 3.5, 3.6, all of those were open and free to use. 3.7 Max is not that. This is their first serious closed proprietary flagship.

They're going after enterprise money now. They want to compete with anthropic and open AI for real business customers. And that's a major strategic move. The model dropped on May the 19th, 2026, quietly one day before the big public announcement at the Cloud Summit.

It's already live on Alibaba Cloud Model Studio, Open Router, Together AI, the platform called Kubrid AI. You can use it right now, pretty much everywhere. And here's where it gets interesting. For anyone running AI agents, Kuen 3.7 Max works across multiple agent frameworks.

So it supports OpenClora, supports Hermes Agent, the self-evolving open-source agent from Noose Research. It supports Cloud Code. It can even support Kuen's own Kuen Code framework. The API is compatible with both the OpenAI spec and the anthropic spec.

So if you're already running OpenClora or Hermes in your business, you can drop Kuen 3.7 Max in as your underlying model without rebuilding anything from scratch, which is actually massive because it means you don't have to choose between the best agent framework and the cheapest frontier model. You can actually just have both. Now let's talk benchmarks because the numbers here are genuinely surprising too. On TerminalBench 2.0, one of the hardest agentic coding tasks out there, Kuen 3.7 Max scores 69.7.

That beats DeepSeek v4 Pro at 67.9. And it beats KimiK 2.6 at 66.7. And it beats Cloud Opus 4.6 Max at 65.4. On SWBench Pro, which tests the model's ability to fix real software bugs, Kuen 3.7 Max scores 60.6.

Cloud Opus 4.6 Max scores 57.3. DeepSeek v4 Pro scores 59. And on MC Atlas, which tests how well a model coordinates with other tools and agents, Kuen 3.7 Max scores 76.4. Cloud Opus 4.6 Max scores 75.8.

On GPQA Diamond, one of the hardest reasoning tests available, harder PhD level science questions, Kuen 3.7 Max scores 92.4. Cloud Opus scores 91.3. So across almost every benchmark that matters for running AI agents in your business, Kuen 3.7 Max is at or above what Anthropx's best publicly available model can do. And the overall artificial analysis intelligence index puts it at 56.6 compared to Cloud Opus 4.7 at 57.3.

0.7 point difference at the frontier level. Now, here's a price comparison. And this is where you need to pay attention. So Cloud Opus 4.7, $5 per million input tokens, $25 per million output tokens.

GPT 5.5, $5 per million input tokens, $30 per million output tokens. Kuen 3.7 Max, 2.5 per million input tokens and 7.5 per million output tokens, right? It's half the input price of Opus and a quarter of the output price. So if you're running heavy AI workflows, agents doing research, writing, analysis, lead generation, the cost difference over a month of usage is real money.

But I have to be honest with you here because there's a catch on the cost side that most people aren't talking about. Kuen 3.7 Max is the boss. So during public and independent benchmark testing, it generated 97 million output tokens. The median for other models in the same test was around 24 million.

So it's outputting about four times more words for the same tasks. At $7.5 per million output tokens, those extra words add up. A task that would cost you 180 with an average model could cost closer to $727 with 3.7 Max depending on what it's doing. So don't just look at the rate card, run your actual tasks for it and measure.

The savings are real, but they're not as dramatic as headline numbers suggest for every use case. There's also a hallucination story here that deserves a closer look. So Kuen 3.7 Max has hallucination rate of 22.9% on the AA Omniscience test. That's the lowest of any frontier model in its comparison group.

That sounds incredible, right? Here's a full picture. The model achieves that low hallucination rate partly by refusing to answer things it's not actually sure about. So its attempt rate, how often it actually tries to answer a question, dropped from 67% on the previous model to 48% on Kuen 3.7 Max.

So it's basically saying, I don't know, more than half the time, rather than really getting it wrong. For certain tasks, that's probably what you want, right? But if you're using it for coding, for analysis, for building workflows, model says, I'm not sure, instead of making something up is a good thing. But if you're building a customer facing chatbot that needs to answer lots of questions with high confidence, the 48% attempt rate is a real problem.

It will leave too many questions unanswered. So you really want to know your use case before you even switch. And speaking of use cases, let me tell you about that 35 hour demo in more detail, because this is genuinely the most ambitious single agent demonstration any major AI lab has published in 2026. So Alibaba gave Kuen 3.7 Max a task, optimize a specific AI computation kernel, a piece of software that handles how AI models process memory.

For Alibaba's own custom chip, the Genwoo M890, this chip has never appeared in any AI training data. The model had never seen it before. No documentation, no example code, just a task description and an evaluation script. Over 35 hours, the model made 32 separate attempts at optimizing the kernel.

It called tools 1,158 times. It diagnosed its own failures, rewrote its approach multiple times, and kept improving even after 30 hours, well past the point where most models would stop making progress. The final result, it was 10 times faster than the original code. So by comparison, GLM 5.1 hit 7.3 times on the same task.

DeepSea V4 Pro hit 3.3 times, and Quantum 3.7 Max more than doubled the second best result. Now for transparency, this is the vendor stated. So Alibaba ran the demo, independent researchers had not reproduced it as of May 25th, 2026. So treat it as a very promising signal, not confirmed fact.

But even if the real world number is half of that, it's still an extraordinary result. And the bigger point is what it shows about where AI agents are heading. Right, a year ago, the idea of an AI agent running for 35 hours and completing a complex engineering task, better than any human team, would have kind of felt like science fiction. In May 2026, Alibaba just did it, on their own hardware, using their own model, with no human in the loop.

That's the direction everything's headed. Now, let me talk about Hermes Agent here as well, because if you're watching this channel, there's a good chance you've already heard me talk about Hermes Agent, the self-evolving agent from news research. It builds skills from experience, runs persistently, has a built-in feedback loop. Quantum 3.7 Max is now listed as a compatible model for Hermes, which means you can use Hermes' self-evolving agent framework, with Alibaba's front-end model, powering it underneath.

Hermes handles the memory, the skill building, the persistence. Quantum 3.6, sorry, 3.7 Max, handles the reasoning, the coding, the analysis. That combination works pretty well. I've tested it out myself, works really well.

And since Quantum 3.7 Max has strong MCP coordination skills, remember 76.4 on MCP Atlas, it should handle multi-tool orchestration well. Now, if you're already inside the Air Profitable Boardroom, you already have access to our AgentOS, the system where you can plug in all your AI agents, including your OpenCore setup, your Hermes agent, and your cloud code workflows into one unified operating system for your business. We've already been testing it with Quentum 3.7 Max integrations inside the agent operating system. And members inside have been asking about this.

So if you want live walkthroughs, a 30-day roadmap for plugging Quentum 3.7 Max into your agent stack, and step-by-step tutorials on setting this up inside AgentOS with coaching calls where you can ask questions live about your specific setup, link in the comments description or go to theairprofitableboardroom.com. Now, let me address something I hear a lot, which is like, Julian, don't need another model. Already used Claude. Why should I even care about this?

Here's the honest answer. You probably don't need to switch everything over today. Claude Opus 4.7 is still slightly ahead on the overall intelligence index. If you have workflows that are working well, don't break them for the sake of it, right?

But here's what you should be thinking about. As AI agents become more central to how your business runs, lead generation, content creation, research, customer follow-up, proposal writing, the cost of running these agents at scale becomes a real business expense, right? $2.5 versus $5 per million input tokens sounds small, but if you're running agents all day, every day, that difference compounds fast. And bear in mind, like we're only using AI agents more and more, right?

We're not using them less anymore. So the businesses that figure out how to run frontier-level intelligence with mid-market prices are gonna have a real cost advantage over the ones paying premium prices for marginally better performance. Coin 3.7 Max is the first time a Chinese model has genuinely closed the gap enough to make that conversation worth having. And this isn't a one-off, right?

The artificial intelligence index shows Alibaba at 56.6, Opus 4.7 at 57.3. Only 0.7 points, right? Six months ago, that gap was much larger. The trajectory's here, right?

And it's clear, which is Alibaba is catching up. There's also a broader thing worth saying as well. We're in 2026. The idea that the AI race is just between Anthropic, OpenAI, and Google was always gonna have a shelf life.

Chinese AI labs have been building hard, right? The Alibaba Cloud Summit framing was, we're building China's AI factory. It's not marketing, right? It's a full-stack strategy.

The own frontier model, the own custom chip, the own agent frameworks, their own deployment infrastructure. This is a complete ecosystem play, not just a model release. If you're building your business on AI right now, you need to understand the models you're using six months from now might look very different from the ones you're actually using today. So what should you actually do with all of this?

If you're running OpenCore, you can configure 3.7 Max as you would through Alibaba Cloud Model Studio or OpenRetail, test it on your actual tasks, content creation, research, lead generation, and compare the output quality and cost against what you're currently running. If you're running Hermes Agent, Quen 3.7 Max is already listed as a compatible model. Plug it in and run your standard Hermes workflows. Through it, pay attention to the verbosity.

Hermes workflows with a lot of back and forth may generate more tokens than expected. And if you're not running any agents yet, this is a great moment to start. The cost of frontier AI level is dropping. The tools for running agents, OpenCore, Hermes, AgentOS are getting more capable and models like Quen 3.7 Max are making it more affordable to run serious AI workflows without enterprise level budgets.

The businesses that win over the next two years are the ones building AI workflows now, while most of their competitors are still doing things manually, which brings me to the biggest limited belief I see holding people back right now. A lot of people think AI is too complicated for you to get set up. And I'm not gonna lie, six months ago, getting AI agents running your business took real technical effort. In 2026, with tools like OpenCore's dashboard, Hermes' built-in tools, and the AgentOS inside the AI platform, the setup has never been simpler, right?

You don't need to touch a line of code. You need to understand what tasks you want automated and follow a step-by-step process for it. Some people as well will say, well, I don't know which model to use. That's actually the wrong question to start with.

Start with the problem you're trying to solve. If it's generating leads, there's a workflow for that. If it's writing content, there's a workflow for that, right? The model choice comes after the workflow design, not before.

And this is exactly the kind of decision we walk through on the weekly coaching course inside the Aircraft Forwarding. Other people say, well, wait until the technology is more stable. Here's the thing. The technology is not going to stop changing.

Quen 3.7 Max just dropped, and it's already competitive with Claude Opus 4.7. In three months, there will be something new. In six months, something else after that, right? The people waiting for stability are the people who will still be waiting in 2027, whilst their competitors have had 12 months of AI automation experience under their belt, right?

The time to learn is now, whilst the tools are still accessible and the competition hasn't caught up yet. And if you want to stay ahead of all this, not just today, but as the models keep improving, come join us inside the Aircraft Forwarding with Bao, the agent operating system, where you can plug in Quen 3.7 Max, OpenCore, Hermes, Claude Code, all your tools into one unified system for your business. We've got tutorials going up, showing exactly how to configure Quen 3.7 Max in your agent stack, how to use it for lead generation and content workflows and coaching calls, where 3,200 business owners are already automating with these tools and can tell you exactly what's working right now. Plus, a 30-day roadmap specifically designed to get your first agent workflows live and generate results.

Link in the comments description or just go to the airprofitable.com. The short version of Quen 3.7 Max is this. It's a Chinese frontier level intelligence, half the price Claude Opus 4.7, works with OpenCore and Hermes agent out of the box and built to run long horizon tasks automatically and autonomously. The caveat is it's verbose, so measure your actual costs, has a high abstention rate on factual questions, so test it on your specific use case and the 35-hour autonomous demo is vendor stated, not independently verified, right?

But the direction is clear. Chinese AI, sorry, China's AI is the frontier now. The price war is real and the businesses that figure out how to use these tools are gonna have a serious advantage over the ones that don't. Thanks for watching.

I'll see you in the next one. Cheers, bye-bye.

More episodes

Browse all episodes →