Julian Goldie breaks down the shocking release of Minimax 2.5, a Chinese AI model that matches Claude Opus 4.6 performance at 95% less cost. Discover how the Forge Framework enables radical efficiency, allowing a 230B parameter model to outperform American giants while bypassing hardware export controls. This episode explores the transition of AI from a luxury to a commodity that is 'too cheap to meter.'
TIMESTAMPS
00:00 Minimax 2.5 Performance vs Claude Opus
01:35 The Efficiency Gap: 10B Active Parameters
02:40 Engineering Intelligence: The Forge Framework
04:15 Cost Breakdown: $0.15 vs $3.00 Per Task
05:50 Minimax Business Growth & Hong Kong IPO
07:15 Bypassing US Chip Export Controls
08:40 The Commodity of Intelligence
10:15 Practical Steps for Founders and Devs
Full transcript
are going to be talking about Minimax 2.5. So a Shanghai startup just dropped an AI model that matches Anthropic's best, and it costs 95% less to run. Let me say that again. Minimax M2.5, released this week, scores 80.2% on SWE Bench Verified.
Claude Opus 4.6, released by Anthropic one week earlier, scores 80.8%. That's a 0.6% gap. And here's the part that should make every executive in Silicon Valley sit up straight. M2.5 costs 1 10th to 1 20th the price.
Dollar per hour continuously at 100 tokens per second. This is the model where frontier intelligence becomes, and I'm quoting Minimax directly here, too cheap to meter. Now, I keep saying this, but the acceleration curve from China is not slowing down. Minimax released M2.1 in January, 2026.
That model scored 74% on SWE Bench. Five weeks later, M2.5 hit 80.2%. That's a six point jump in a month. For context, it took American Labs the better part of a year to make that kind of leap.
And Minimax did it with a model that's only 230 billion parameters total, 10 billion active at any given time. That's the thing that keeps getting lost in these conversations. They're not just catching up on performance, they're doing it with radically more efficient architectures. And that's how powerful this stuff is, all right?
So how did they do it? Let me break down the one technical detail that actually matters here, all right? And Minimax actually built something called Forge Framework. Think of it like this.
Forge Framework is what most, well, most AI models learn by being graded on their answers after the fact, right? You ask it a question, a human says, good job or bad job, and the model adjusts. Forge is different. So it drops the AI into 200,000 real-world coding environments and lets it learn by doing.
And so this is what's interesting here. It's like actual GitHub repositories, actual bugs, actual multi-file code bases across Python, Rust, Go, TypeScript, C++, and seven other programming languages. It's the difference between studying for a test and doing a two-year apprenticeship. The gap is huge.
And the result is a model that doesn't just answer coding questions, it thinks like a software architect. Before it writes a single line of code, M2.5 plans to structure, maps out the features, designs the interface. That behavior wasn't programmed in. It emerged during training.
On the multi-SWE bench test, which measures performance across multiple code bases simultaneously, M2.5 actually beats Opus 4.6, 51.3% versus 50.3%. On function calling, the ability to chain together multiple tools in sequence, which is what matters for real autonomous agents, M2.5 scores 76.8 on the Berkeley function calling leaderboard. That's higher than Claude, higher than GPT 5.2, higher than Gemini 3 Pro. And it completes SWE bench tasks in 22.8 minutes on average.
Claude Opus 4.6 takes 22.9 minutes. Virtually identical speed, but here's a kicker. The cost per task on M2.5 is roughly 15 cents. On Claude Opus, the cost per task is about $3.
That's a 20X difference. So who's behind it? Well, Minimax was founded in December, 2021 by Yan Yunji. I've probably totally decimated that name, but Yan Yunji is the way that I pronounce it.
A former vice president at SenseTime, one of China's biggest AI companies. Now they raised $600 million led by Alibaba in early 2024. At a $2.5 billion valuation. Then in January, 2026, they went public on the Hong Kong Stock Exchange.
IPO was oversubscribed 1,837 times. Let that number sink in. Nearly 2,000 times oversubscribed. Shares doubled on the first day of trading, closing up 109%, pushing the market cap past 100 billion Hong Kong dollars.
Roughly $13.7 billion US dollars. The stock jumped another 15.7% the day M2.5 was announced. And they're not just a research lab. Minimax has over 212 million users across more than 200 countries.
They built Talkie, an AI companion app, by mid-2024 was the number one AI companion app in US downloads, beating Character AI. More than half its 11 million monthly active users were American. They beat Halo World, right? Which is a video generation platform, pulling in millions of monthly visits.
Revenue went from 3.5 USD in 2023 to $30 million in 2024, to $53 million in just the first nine months of 2025, with 70% of that revenue coming from outside China. This is a company of real products, real users, and real commercial traction. And that's what's one of the most interesting things here. So here's where it gets really interesting from a geopolitical standpoint.
The United States has spent the last three years trying to kneecap Chinese AI through chip export controls, cutting off access to NVIDIA's best GPUs. And what happened? Minimax built a frontier model with only 10 billion active parameters and matches. Models trained on clusters American labs spent hundreds of millions of dollars assembling.
On the VentureBeat podcast, First AI, Minimax engineer Olive Song said the model was trained over just two months using their Forge reinforcement learning framework. They're not brute forcing their way to intelligence with more compute. They're engineering their way there with better algorithms. And they open source the whole thing.
Full model weights on Hugging Face under a modified MIT license, you can download it right now. You can deploy it on your own servers with VLM or SG Lang. Minimax are using M2.5 internally. 30% of all tasks at Minimax headquarters are actually completed by the model.
And 80% of their newly committed code is generated by M2.5. They're not just selling this, they're actually running their company on it. Now, I get the skepticism. Benchmarks aren't everything.
Self-reported numbers always deserve scrutiny. And Minimax is still burning cash, right? $512 million in losses in the first nine months of 2025 alone. They're in what they call a nascent stage of monetization.
So yes, there's a gap between benchmark performance and sustained profitability. I also get the people are tired of hearing everything is changing. Every couple of weeks there's a new update, right? But here's what I can't ignore.
Open Hands team, which builds open source coding agents, got early access to M2.5 and independently tested it. Their conclusion is the first open-weight model to match Claude Sonnet's quality. They called it a two-horse race now. Claude Opus on the premium end, M2.5 on the cost-efficient end.
At a price point roughly 13 times cheaper than Opus. The Kilo Code team tested it and found it achieving 80.2% on SWE bench out of the box, whilst Opus needs specific prompt modifications to push past 80. Let me zoom out for a second. What does this actually mean for people who aren't AI engineers?
It means the cost floor just dropped through the basement. If you're a startup founder, you can now run frontier-level AI agents continuously for the cost of a Netflix subscription. Four M2.5 instances running 24-7 for a year costs under $10,000. So that's a team of AI workers coding, searching, analyzing documents, building PowerPoints, doing financial modeling in Excel for less than a single junior developer's monthly salary in most major cities.
If you're a manager at a mid-sized company, this is the moment to start piloting, not next quarter. Now, because your competitors in Asia are already deploying this. If you're a developer, M2.5 already works in Cursor, Client, Kilo Code, Claude Code, and Overhands. You can swap it in today.
It's free on several platforms for a limited time. There's literally no reason not to try it this week. And if you're in policy or leadership, you need to understand something fundamental. Export controls didn't stop Chinese AI from reaching the frontier.
They may have slowed it down by a few months, but what they actually did was force Chinese labs to become radically more efficient. Minimax is proof of that. The model was smaller, cheaper, faster, and within a rounding error of the best models on Earth. That's not a failure of the controls, it's just reality and policy needs to account for that reality.
The deeper tension underneath all of this isn't really about China versus America, it's about what happens when intelligence becomes a commodity. When the best AI in the world costs a dollar an hour. The bottleneck isn't technology anymore, right? It's imagination.
It's knowing what to build, it's knowing how to manage a hybrid teams of humans and AI agents. Most importantly, it's knowing how to take care of each other as the ground shifts under all of our feet. Because I guarantee you, there is someone watching this right now whose job looks very different from a year from now. Not because they're going to be replaced by a robot, but because the tools available to them and to their competitors just fundamentally changed.
And the question isn't whether to engage with this, the question is whether you do it on your terms or someone else's. Stay sharp and I will see you on the next one. Thanks for listening. Just to note here as well, if you haven't already, check out the AI Profit Boardroom.
This is my AI automation community where we show you practical ways to save time, scale your business and get more customers using Dice. Feel free to check that out. It's available at AIprofitboardroom.com. AIprofitboardroom.com.
More episodes