Minimax M2.7: This AI Just Improved Itself 100 Times
Discover how Minimax M2.7 achieved a breakthrough in self-evolution by optimizing its own code over 100 rounds to become 30% more efficient without human help. We break down how this ultra-low-cost model challenges industry leaders like Claude Opus and why autonomous agent loops are the future of work.
00:00 - Intro: The AI That Fixes Itself
01:20 - Who is Minimax?
02:03 - Defining Self-Evolution
03:12 - Benchmarks vs. Claude Opus
04:27 - 1/20th the Cost of Western AI
05:24 - Native Agent Teams & Skills
06:48 - OpenCore & Scaling Output
08:14 - 4 Steps to Start Today
Full transcript
Minimax M2.7, China built AI that improves itself. So an AI just spent 100 rounds fixing itself. Not a human tweaking it, not a team of engineers writing new code. The AI looked at what it got wrong, planned what to change, rewrote its own instructions, ran the test, checked the results, and did it again, 100 times in a row.
And then when it was done, it was 30% better than when it started. That happened recently, right? It just dropped from Minimax, a Shanghai-based AI company, tasked an internal version of their new model, M2.7, with optimizing its own performance on internal scaffold. The model ran entirely on its own, implementing a loop of analyzing failure trajectories, planning changes, modifying scaffold code, running evaluations, comparing results, and deciding whether to keep or revert every single change for over 100 rounds.
The result was a 30% improvement on performance, discovered entirely by the model itself. It found optimal sampling parameter combinations, more specific workflow guidance, and loop detection optimizations on its own. Nobody told it what to look for. This dropped March 18th, 2026, from a company most people in the West have never heard of.
And if you're thinking, okay, that's just coding stuff, it doesn't affect me, stay with me, because what I'm about to talk about through, and what I'll talk you through today, changes things for everyone who uses a computer to do their job. So first, who is Minimax? Minimax is a Shanghai-based AI company founded in 2021 by Yan Yunji, a former deputy director of Huawei's AI lab. The company raised $619 million in funding, and IPO'd on the Hong Kong Stock Exchange in January, 2026, at a $6.5 billion valuation.
Investors included Alibaba, Tencent, and Goldman Sachs. These are not hobbyists. These are world-class researchers, well-funded, moving fast. They released Minimax M2 in October, 2005, then M2.1, then M2.5 in February, 2026, and now M2.7.
Four significant releases in five months. Most Western labs take six to 12 months between major releases. Let me explain what self-evolution actually means, because the phrase gets thrown around a lot. In the old world of AI, here is how improvement works.
You have a model. Humans watch what it does. Humans figure out what went wrong. Humans write new training data, new rules, new code.
Then they train a new version. The model had no say in any of it. Previous AI updates were pushed forward by researchers, designing experiments, tuning parameters, running tests. The model was just an executor, not involved in the iteration process.
M2.7 changes that. During its own development, M2.7 built its own agent harness, described and designed its own skills, ran its own experiments, analyzed results, and then optimized its own testing framework. The AI is now writing its own performance reviews, finding its own gaps, deciding what to fix, actually fixing it. Another skeptic response here is, it's not really thinking itself, right?
It's just running a script, and that's fair. This is not conscious self-awareness, but here's why that objection misses the point. Self-evolution in this context means the AI is closing the feedback loop that used to require human researchers. Less time between model has a weakness and model fixes a weakness, the improvement cycle therefore gets faster without human labor, and that's what matters.
Now the numbers. According to Minimax's internal benchmarks, M2.7 handled between 30% and 50%, resulting development workflow autonomously, including monitoring experiments, analyzing logs, and implementing code repairs, 30% to 50%, handled without a human touching it. On real-world benchmarks, M2.7 scored 56.2% on SWE Pro, 57% on Terminal Bench 2, and achieved a 1495 Elo on GPT-VAL-AA, setting a new standard for multi-agent systems operating in real-world digital workflows. SWE Pro is not an easy test.
It tests whether an AI can solve actual real programming problems from real GitHub repositories, real bugs, real messy code, real deadline pressure. M2.7 scored 56.22% on SWE Pro, placing it within striking distance of the industry-leading Opus models from Claude. Striking distance of the best models in the world from a lab most people learned about this week. And on the office work, most of you actually do every day, M2.7 achieves an Elo score of 1495 on GPT-VAL-AA, the highest among open-source models with significant improvements in complex editing for Word, Excel, and PowerPoint.
And here's the number I think matters most, the price. Minimax M2.7 scored 30 cents per million input tokens and $1.20 per million output tokens. Compare that to the alternatives. A workload of 10 million input tokens and 2 million output tokens per day costs roughly $4.78 per day with Minimax versus $100 a day with Claude Opus 4.6.
Over a month, that is $141 versus $3,000. Over a year, $51,000 versus a million dollars. Same workload, 51,000 versus a million dollars per year. The expensive models used to be the only good ones.
It stopped being true a few months ago. And most people haven't caught up to that yet. That gap between what's possible right now and what most people think is possible, that's what's gonna separate the people who win in this next period from the people who fall behind. Now, what does M2.7 actually do in the real world?
Well, Minimax M2.7 is trained for production-grade performance on workflows like live debugging, root cause analysis, financial modeling, and full document generation across Word, Excel, and PowerPoint. These are not edge use cases. These are the everyday jobs of analysts, engineers, consultants, finance teams, and marketing teams. M2.7 introduces native support for agent teams, clusters of AI agents that collaborate with distinct roles, allowing agents to challenge each other's logic and conduct adversarial reasoning to reduce errors.
Multiple agents with different roles checking each other's work. So one could research, one could write, one could edit, one could fact check, run at the same time at a fraction of what one human hour costs. On 40 complex skills requiring over 2,000 tokens, M2.7 maintains a 97% skill adherence rate. That's a model you can actually rely on to follow complex, multi-step instructions and still be on the right track at the end.
And this is the right moment to talk about the AI Profit Boarding because what I just described, the agent setup, the model selection, the workflow automation, the specific tools and how to configure them, that's exactly what we break down inside the community. Not in theory, step-by-step with real examples and real workflows you can deploy this week. If you're a business owner, creator, freelancer, consultant, the people inside the AI Profit Boarding are already three steps ahead of where most businesses are right now. Go to the AIProfitBoarding.com or link in the comments description to join us.
Now, let me connect M2.7 to OpenCourt because this combination is important. OpenCourt is an open-source AI agent operating system that turns any computer into persistent multi-platform AI agents. It supports isolated agent workspaces, tool core, image processing, and routing across messaging apps, handling coding agents, personal assistants, and multi-agent teams with seamless model swapping. M2.7 was built to run inside OpenCourt natively.
Community feedback confirms M2.7 outperforms M2.5 in OpenCourt on complex instruction following, multi-turn debugging, and long-term context persistence. And it connects to messaging apps like Telegram, WhatsApp, Discord, so your AI agent lives in the app you already use. Now, two companies. You've got company A, which has 10 employees that doesn't use AI agents.
You've got company B, there's 10 employees, each running two or three AI agents, handling background tasks constantly. Company B just scaled to 30 employees worth of output without hiring anyone new. Now, at the price, company B is running frontier level AI for a 20th of what it would have cost a year ago. Their cost per output unit can collapse, right?
Now, add the self-improvement loop. The AI agents company B uses are getting better over time on their own, not just from new model releases, from the iterative improvement loop built into the model itself. Company A is competing against a team that is effectively three times the size, costs a fraction per unit of output, and is getting better automatically whilst they sleep. It's not a fair fight.
And the crazy part is, most businesses are still company A right now. Here are four concrete things you can do starting today. If you're completely new, start with the concept of AI agents, not just AI chat. The difference between talking to AI and having AI work for you is enormous.
Find one repetitive task you do every week that takes two or more hours. That's your starting point. If you're already using AI for tasks, move from single turn to multi-agent. Instead of asking one question at a time, build workflows where research gets handed to drafting, drafting to editing, editing to publishing.
One workflow starts small. You can do that with Minimax M2.7. And if you're building with AI already, test M2.7 now. It's available through the Minimax API at 30 cents per million input tokens and $1.20 per million output tokens.
Run it against your current model on your actual use case. Given the price difference, even if it's 90% as good, the economics might be dramatically better. And if you lead a team or run a business, stop asking, how do I add AI to what we're doing? And start asking, which jobs does my team do today that could be handled by a well-configured agent like Minimax?
The AI handles the first 70 to 80%. Your people focus on the last 20 to 30% that requires real human judgment. And that's the model that wins. Minimax's move toward recursive self-improvement signals a shift in the industry, a future where the models we use are as much the architects of their progress as they are the products of human research.
Here's what that means for Pace. We've been measuring AI progress by releases. Each one takes months. And that model of progress assumes human research is always a bottleneck.
The AI has to wait for humans to figure out what went wrong. What happens when the AI starts closing its own improvement loops? Well, the Pace doesn't just stay the same, it accelerates. Because now the AI itself is contributing to getting better, running 24 hours a day, running 100 iterations in the time a human team could run one.
The people not using AI agents think the slope looks slow. The people running agents, using tools like OpenGL or models like M2.7, well, they're watching the cliff approach, right? Self-improvement is a moment you go from slope to cliff. So here's the headline one more time.
Minimax tasked M2.7 with optimizing its own performance. The AI analyzed its failures, planned its changes, modified its own code, ran its own evaluations, and repeated that cycle over 100 rounds. It finished 30% better than when it started without a human telling it what to fix. That is a new thing in the world.
It happened yesterday from a company that's released four major models in five months at a price point that is 1 20th of the Western alternatives. The self-improvement loop is open. The question is whether you're going to be inside it or watching from the outside. Go to the AI Profit Boarding, link in the comments description, or go to the AIProfitBoarding.com.
You can connect with me personally, check out our community, learn how to implement this stuff inside your business. And it's exactly where we show you exactly how to step through this self-improvement loop. See you in the next one. Thanks for watching.
More episodes