AI News Today
← All episodes
Episode 1 · March 19, 2026 · 21:10

Xiaomi's Secret AI Is FREE & Beating Claude

Xiaomi’s Secret AI: Mimo V2 Pro Is Beating Claude for Free


Xiaomi has shocked the tech world with Mimo V2 Pro, a trillion-parameter AI model that matches top-tier performance at a fraction of the cost. Learn how this 'quiet ambush' is disrupting the industry and what it means for the future of autonomous AI agents.


00:00 - Intro: The Mystery of Hunter Alpha

01:52 - Xiaomi’s Secret AI Revealed

02:25 - Specs: 1 Trillion Parameters

04:43 - Benchmarks & Cost Comparison

08:30 - Multi-Modal & Voice Releases

11:11 - Real-World Agent Capabilities

13:00 - System-Level Phone Integration

18:07 - 3 Ways AI is Changing Forever

Full transcript

Xiaomi's secret AI is free and beating Claude. So a phone company just built an AI that nobody could identify. For eight straight days, developers all over the world were using a mystery model. It had no name, no company behind it, just a code name, Hunter Alpha.

And it was beating models that cost 10 times more to run. It topped the charts on Open Router. That's the world's biggest platform where developers plug their apps into AI models. And Hunter Alpha sat at number one day after day, processing 500 billion tokens.

A week, making it the platform's highest performing model. Just to put that number in plain English, a token is roughly three quarters of a word. 500 billion tokens a week means hundreds of billions of words being sent through this mystery model by developers who had no idea who made it. That's not a niche experiment.

That's a model quietly becoming part of real production systems at scale. The Internet started going crazy. Everyone thought it was DeepSeek. DeepSeek is a Chinese company that shocked the world earlier this year.

Their cheap, powerful models sent tech stocks crashing because suddenly investors were asking, do you actually need to spend billions of dollars to build world class AI? When Hunter Alpha showed up with specs that rumored DeepSeek v4, a trillion parameter, one million token context window, people assumed the DeepSeeker quietly launched the next thing ahead of schedule. Reuters ran the story. Polymarket even opened a betting market on it.

The model itself, when asked who built it, gave the most mysterious answer possible. Only know my name, my parameter scale and my context window length. It said that's it, an AI that refused to name its own creator. During Reuters testing, the chatbot itself introduced itself as a Chinese AI model, primarily trained in Chinese, and said its training data reached May 2025.

The exact same knowledge cut off the DeepSeek's own chatbot reported. That detail fuel on the speculation. Every point, everyone pointed at DeepSeek and then the reveal happened. It wasn't DeepSeek.

The model was revealed to be from Xiaomi. The Chinese smartphone and EV giant. Xiaomi, the company that makes earbuds and budget phones and electric cars. That's Xiaomi.

And here's what makes this more than a fun mystery. The model was free. It matches Claude Sonnet and Kodi. It almost catches Claude Opus on agent tasks.

And it costs one fifth of what Opus charges per token. Now it's a phone company walking into a room full of the most well-funded AI labs on earth and saying, we built something that completes with your best work and we're giving it away for free. So let's talk about what this model actually is before we get into what it can do. MIMO V2 Pro has a total number of parameters exceeding one trillion.

The number of activated parameters is 42 billion, which is about three times that of the previous generation model. It supports a context length of one million tokens. A trillion parameters means this model has a trillion little dials and switches inside it, all trained to understand language, write code and reason through problems. The more parameters, the more the model can hold in its head.

Claude Opus 4.6 operates at a similar scale. That's not a small demo or a model. That is a heavyweight, right? And the one million token context window means you can feed this model about 800,000 words in a single conversation.

That's roughly eight full-length novels at once. You could paste in an entire code base, years of accumulated code from an entire engineering team. And the model holds all of it simultaneously. Most older models topped out at 4,000 tokens, about three pages of text.

The jump from 4,000 to a million is not a small upgrade. That's a completely different class of capability. Now, the 42 billion active parameters matters because you don't activate the full trillion every time someone sends a message. The model only fires up the portion it needs for a given task.

Think of it like a company with a thousand employees, where only 42 show up for each specific meeting, depending on what that meeting is about. It keeps the model fast and affordable to run without sacrificing the depth that comes from having a trillion parameters available on reserve. MIMO V2 Pro uses a hybrid attention mechanism with a seven to one ratio, up from five to one in the previous flash version. This lets the model process massive context efficiently without sacrificing speed.

So here's a plain English version of that. Imagine you're reading an 800 page book and you need to remember everything to answer questions about it, right? Full attention means reading every word with complete focus. Efficient, but slow.

Sliding window attention means reading in chunks, right? So paying close attention to what's right in front of you and checking back on earlier parts only when necessary. The hybrid approach mixes both. For most of the document, you use a fast, efficient method.

For the parts that matter most, you use full focus. Seven to one means for every seven sanctions you read efficiently, one gets full attention. Bumping that ratio up from five to seven is what lets this model handle a million tokens at a cost that doesn't make it economically pointless. Now, the benchmarks, because this is where the story gets interesting.

On Claude Eval, a benchmark for agentic scaffolds, the model scored 61.5, approaching the performance of Claude Opus, 4.6 to 66.3, and significantly outpacing GPT 5.2 at 50. Claude Eval measures something specific, not general intelligence, not how well the model can write an essay or explain a concept. It measures how well a model can act as an autonomous agent. It uses tools, calls external APIs, plans multi-step tasks, implements them without a human confirming every single move.

That's the ability that actually matters for building serious AI systems like OpenClaw in 2026. And Maimo V2 Pro is scoring 61.5. Claude Opus, 4.6, Anthropics' best model, the most capable version of the AI system that's widely considered the industry leader, scored 66.3. So we're talking about a gap of just under five points from a team that's been building AI for two years.

On coding specifically, it's even sharper. On pure coding tasks, Maimo V2 Pro outperforms Claude Sonic, 4.6, on SWE Bench Verified. SWE Bench Verified is one of the most respected coding benchmarks in the industry. It uses real GitHub issues, right, so actual bugs and problems from actual software repositories.

You give the model the issue, it writes a patch to fix it, and the patch driver passes the unit tests, or it doesn't. No partial credit, no benefit for doubt. It works or it doesn't. And Claude Sonic, 4.6, scores 79.6%.

Maimo V2 Pro beats that number from a team that didn't exist even two years ago. And then there's a cost. Sorry, Artificial Analysis reported that running their full evaluation suite costs only $348 for Maimo V2 Pro compared to $2,304 for GPT 5.2 and $2,486 for Claude Opus 4.6. The same tests, the same tasks, seven times cheaper for nearly the same result.

The pricing of Maimo V2 Pro's API is only one fifth of that of Claude Opus 4.6. And right now, this week, it's completely free. Maimo V2 Pro is partnering with OpenCore, OpenCode, KiloCode, Blackbox and Klein to offer one week of free API access for developers worldwide. So you don't need to subscribe, you don't need to pay anything.

If you build things with AI, apps, automations and agent systems, you can test this model at zero cost today. Now, let me tell you about the person who actually built this. Xiaomi's AI model team, Maimo, is led by Luo Fuli, a former DeepSeek researcher. She's been called a genius in the Chinese tech industry.

She came directly from the team that built the models that sent global tech stocks into a spin at the start of this year, sorry, at the start of last year. She left and joined Xiaomi's AI division, which was founded in 2024, two years old. Less than that, actually, when you account for how long serious model development takes. She called the launch a quiet ambush, not a big press conference, not a carefully staged reveal event.

A model that showed up anonymously on a public platform, let developers use it for eight days without knowing who built it, topped the charts based purely on performance and then revealed itself when it was ready. That is a specific kind of confidence, right? You don't do a blind test unless you're certain what the result will be. Luo Fuli stated that the company plans to open source a model variant from this latest release when the models are stable enough to deserve it.

When that happens, and it will happen, developers everywhere can run this model locally with no API costs, no usage limits, no subscription, just a trillion dollar parameter model on someone's server, building whatever they want. DeepSeek did this with v3 and r1, and it fundamentally changed what was available to the open source community overnight. MIMO is doing the same thing and adding another serious option to that ecosystem. And it wasn't just one model, right?

In the early morning of March 19th, Xiaomi released the flagship base model MIMOv2 Pro, the fully modality agent model MIMOv2 Omni, and the speech synthesis model MIMOv2 TTS. MIMOv2 Pro is the coding agent and agent brain we've been talking about, but MIMOv2 Omni is a different animal. This is positioned as Xiaomi's fully modal foundational model, integrating multimodal comprehension capabilities for images, video and audio, along with powerful agent abilities, right? According to the Goldman Sachs report, this model matches or surpasses Gemini 3 Pro, Quad Opus 4.6, Gemini 3 and GPT 5.2 across key metrics in audio, image, video understanding, agent capabilities.

It matches or surpasses Quad Opus 4.6 on multimodal tasks from a company that most people 18 months ago associated with reasonably priced smartphones. And then MIMOv2 TTS voice model, it aims to enable intelligent agents to communicate with people using a warm, emotional and soulful voice with highly controllable, multi granular style control and natural rhythm reproduction. Three releases in one day, a coding and agent brain, a model that can see, hear and watch video, and a voice synthesis model so your agents can talk back. That's a full stack.

It's not a company dabbling in AI. That's a company making a serious, sustained push to build the complete infrastructure for the agent era. Now, let me handle the skepticism directly because it's legitimate. Benchmarks are one thing, real world performance is another.

Companies publish numbers and make their models look great all the time. Half the benchmarks out there are games. So how do we actually know this works? The Hunter Alpha test was a real answer to that question.

Eight days, real developers, real production workflows, real coding tasks, and crucially, nobody knew who made the model. There was no brand association, no reason to be generous, no incentive to report good results. They used what was available and it rose to the top organically. During the Hunter Alpha test phase, the top apps by core volume were all coding focused tools, confirming MIMOv2 Pro's high usability and reliability in real development workflows.

That is the cleanest possible signal. Developers voted with actual usage, blind over eight days, a trillion tokens worth of voting. And Artificial Analysis, an independent third party benchmarking organization, verified the numbers themselves, not just the claims on Xiaomi's website. They placed MIMOv2 Pro at number 10 on the Global Intelligence Index with a score of 49, in the same tier as GPT 5.2 Codex and the head of Grok 4.2 Beta.

So this isn't Xiaomi grading their own homework. Now I want to show you what agents can actually do with this model because a benchmark score tells you one thing. Watching it work tells you something more visceral. In intelligent agent frameworks such as OpenCore and CoreCode, MIMOv2 Pro can complete complex workflow orchestration, long-term planning, and precise tool invocation without manual intervention.

An agent is not a chatbot. A chatbot waits for you to ask it something and it answers. That's it. An agent goes out and actually does things.

So it can write a script, run the script, check whether it works, fixes it when it breaks, and then run it again. Can search the Internet, read what it finds, decide what to do with that information, and take action. No confirmation required at any step. No human in the loop for every decision.

And Xiaomi proved this worked with a live demo during the launch. A tester asked MIMOclaw to design a website that updates the list of companies listed on the Hong Kong Stock Exchange and A-shares every day at 7pm. MIMOclaw used a Python crawler to regularly fetch data and generate a static page for direct development. After detecting mismatches during the running test, it corrected and supplemented the data automatically.

Someone typed what they wanted, they write the code to build it, deployed it, ran it, found an error, fixed the error, delivered a working live product, and there was no coding knowledge needed from the human. No back and forth troubleshooting, no stack overflow. Just describe the outcome and the agent figures out how to get there. That is where we're at right now today with a model that's free.

Goldman Sachs mapped out the roadmap. MIMOv2 Pro's next target is high complexity reasoning and long-term task planning. MIMOv2 Omni aims for cross-hour and even more cross-day continuous intent planning, real-time stream sensing, and execution of robot and hand actions. Let me tell you what that means in plain English, right?

Cross-day continuous intent planning. That basically means an agent that doesn't complete a task in a single session forgets about it. It's an agent that holds a goal across multiple days, monitors conditions, takes actions when the right moment arrives, adjusts the strategy when things change, right? So you give it a business objective and it works toward it autonomously.

It's not a chatbot. That's something that behaves more like a junior employer than a calculator, right? And here's where the China piece of this matters. So smartphone makers like, for example, Xiaomi, Honor, and Huawei are scrambling to integrate system-level agents into their devices, capitalizing on the open source agent framework, OpenClaw, right?

So OpenClaw is a framework that's been sweeping through the Chinese tech industry in 2026. It lets a model take real actions in the real world. In Tencent Launch, for example, Tencent Cloud launched a one-click deployment template for OpenClaw. The number of cloud users deploying it has exceeded 100,000 and continues to grow.

Basically, every major Chinese tech company is racing to build on top of OpenClaw. Xiaomi's MyClaw system integrates more than 50 functions into the phone, including reading and sending messages, managing calendar events, searching the internet and launching apps and accessing files. That's 50 functions baked into the operating system. Not a separate app where you remember, you know, use it when you want to remember it.

It's an AI woven into the phone itself that can see your calendar, read your messages, launch your apps, execute tasks whilst you're doing something else entirely, right? You ask your phone to plan your week and it reads your calendar, checks your emails and looks up relevant deadlines and adds tasks directly. Currently in closed beta on the Xiaomi 17 series, but the public version is coming soon. And all three models have been integrated into WPS Office, right?

Xiaomi phones and also computers via the MyClaw agent system and the Xiaomi browser, right? So this AI doesn't just live in the cloud as an API call. It lives in the hardware that hundreds of millions of people already use every day. That's a physical AI strategy.

And that distinction between AI as a cloud service and AI as something embedded in every device you own, it's going to matter a lot over the next two years. So let me zoom out to the bigger pattern, because this doesn't make full sense in isolation. For years, the race to build powerful AI was understood to be a contest between a handful of tech giants. Anthropic, OpenAI, Google, DeepSeek, ByteDance and Alibaba with bottomless budgets and decades of institutional knowledge.

That assumption is breaking down fast. DeepSeek showed it first. A relatively small team with efficient training methods, built models that genuinely competed with OpenAI's best. People said it was a one-time thing.

Then Minimax showed up. Then Alibaba's Quen series. Then Kimi, now Xiaomi's Maimo. Two years old, led by one former DeepSeek researcher, releases a model that outperforms Claude Sonnet on coding and sits within five points of Claude Opus on agents.

The 25x price gap between the cheapest and most expensive front-end model is the biggest change in 2026, right? It's 25 times. The same capability tier, similar benchmark performance and a 25x difference in what different providers charge for access. So in December 2025, frontier coding required Opus tier pricing of $5 input and $25 output per million tokens.

In March 2026, Minimax M2.5, for example, delivers 80.2% SW bench at $0.30 input and $1.20 output. So the price collapse is real and it's accelerating. More teams competing at the frontier means faster improvement and lower prices simultaneously. That combination, better and cheaper at the same time, doesn't happen often in technology, but it's happening right now in AI every single quarter.

The frontier is flattening. The question now is not whether AI's industry in China catches up. It's how the incumbents respond when it does. For the business owners and creators watching this, that competitive pressure is your advantage.

Every time serious teams enter the race and new ones do all the time, your cost of accessing top tier AI drops. Every time a model like MIMO V2 Pro launches free, the capability that was behind a paywall yesterday is now available to anyone with an internet connection. And this is where I want to be direct with you, right, about what this actually means. Every time a new model drops at a lower price with higher capability, the people already know how to use AI agents, who know how to build automations, how to set up workflows, how to plug these models into real systems.

They get to take advantage immediately, right? They see MIMO V2 Pro and they think, great, that's a cheaper brain for my agent stack. Let me swap it in and cut my API costs by 18% this month. The people who don't understand any of this, they just keep watching, the gap keeps growing, right?

And the gap isn't just about awareness anymore, it's about compounding. The businesses that started using AI agents seriously six months ago have now thousands of hours of tasks running through those systems, right? They've refined their prompts, they've built libraries of automations, they've freed up time that they've reinvested into building more. The gap between them and someone starting today isn't six months, it's a compound value of six months of learning and implementation.

That's exactly why the AI Profit Boardroom exists. It's where I work directly with business owners, creators, and freelancers to actually implement this stuff. Not just understand it conceptually, but build real systems that save time and generate real results. And every major model release, every new tool, every practical use case, we go deep on how to use it in your actual business, not just describe it.

Go to the AIprofitboardroom.com or link in the comments and description to join us. So let me give you three things this story actually tells us about where AI is going. First, the geography of AI development has permanently changed. Frontier AI used to mean a handful of companies in San Francisco and Seattle.

That's no longer true. The Maimo team is in Beijing. Their lead researcher came from Hangzhou. China's massive AI talent density, so many Chinese companies can now catch up, right?

Because they've got so much talent in that industry. And that's a sentiment that's recurred across dozens of independent testers during the Hunter Alpha period. For you as a business owner or creator, it's not a threat. It's an opportunity.

More serious teams, competing means more options, lower prices, better tools. The model power in your next automation might come from Xiaomi or a Chinese startup or someone you've never heard of. And it will work just as well as something you're paying a premium for today. Second, the agent layer is where the real race is happening, right?

Xiaomi didn't optimize Maimo V2 Pro to be a better chatbot. They built it specifically for open core, for core code and for agent scaffolds. The benchmarks they led were with agent benchmarks. They are building infrastructure for autonomous AI workers, not better assistants.

The companies and individuals who understand agents, who know how to give these models real tasks, real tools, real authority to go and implement, are going to move dramatically faster than everyone else. A well-built agent doesn't just save you 30 minutes a day. It handles entire workflows whilst you sleep. Third, open source is the long game and it keeps working, right?

Keeps winning. When Xiaomi releases the open source version of this model, it joins DeepSeek V3, DeepSeek R1, and a growing list of serious open weight models that anyone can run, modify and deploy at zero cost. The best models on earth are becoming free, not eventually now. So here's the close.

The Hunter Alpha story is fun. A mystery model, a reveal, a phone company embarrassing the AI incumbents with their own benchmarks. But the story underneath is more important than the drama. The benchmarks are converging.

The prices are collapsing. The teams doing it are, in some cases, less than two years old. The tools for doing serious AI work, real-gentic work, not just chatting, just got dramatically more accessible. A small business with a limited AI budget is now playing in a different league than they were six months ago.

The things that were economically out of reach are free this week. And the question is no longer like, can it just talk? But can it act? A Maimo V2 Pro can act.

The benchmark scores within five points of the best models in the world, at one fifth of the cost of the most expensive ones. Free right now, today, through OpenRouter and five major Asian frameworks. You don't have to be a developer to care about that. You just have to understand that the tools for doing serious automated work with AI just got cheaper and more capable at the same time.

And the people who figured out how to use these tools first are the ones who are going to look back in two years and understand exactly why today mattered. Don't be the person who waited, and I'll see you in the next one, or feel free to check out the AI Profit Boardroom. Link in the comments description. You can connect with me personally.

You can check out all our 30-day roadmaps and the prompts we put inside there. You can see new daily tutorials on all this stuff, plus how to use OpenCore, and I hope to see you inside. Cheers.

More episodes

Browse all episodes →