AI News Today
← All episodes
Episode 1 · February 12, 2026 · 09:57

New GLM 5 is POWERFUL!

Julian Goldie breaks down the surprise release of GLM5, the Chinese open-source AI model outperforming top US models. Built entirely on Huawei silicon, it bypasses US export controls using Mixture of Experts and SLIME architecture.


TIMESTAMPS

00:00 The Mystery of Pony Alpha

01:15 What is GLM5?

02:30 Breaking the Nvidia Dependency

04:00 Mixture of Experts & SLIME

05:45 Agent Mode & Autonomous Tasks

07:00 Pricing vs Claude & GPT

09:00 The Geopolitical Impact

Full transcript

A mystery model just showed up on the internet. No announcement, no press conference, no company name attached. Just a codename PonyAlpha and an API endpoint routing traffic straight to Beijing. Within 24 hours, it processed 40 billion tokens and 206,000 requests.

Developers were losing their minds. People thought it was a secret clawed release. Others thought DeepSeek dropped a stealth v4. Somebody even guessed it was Grok, and it wasn't any of them.

It was GLM5 from Jerpoo AI, a company most of you have never heard of. And it just became the number one open source AI model on the planet. Let me give you the numbers so you understand what we're talking about here. 744 billion parameters, only 40 billion active at any given time.

Trained on 28.5 trillion tokens of data. A 200,000 token context window. And a maximum output of 131,000 tokens, which is one of the highest in the entire industry. On SWE Bench Verified, the gold standard coding benchmark, GLM5 scored 77.8%.

For context, Claude Opus scored 80.9%. Gemini 3 scored 76.2%. This open source model from China is sitting right between Google and Anthropics best. On BrowseComp, which tests web retrieval and information synthesis, GLM5 scored 75.9.

Number one among all models tested. Not number one open source, number one period. And here's the part that should stop you in your tracks. It was trained entirely on Huawei Ascend chips.

Zero NVIDIA GPUs, zero American-made hardware, not a single chip touched by US export controls. They used Huawei's MindSport framework and built the entire training pipeline on domestic Chinese silicon. And I keep saying this because I need to hear people to hear it, right? The assumption that cutting off chip exports would slow down Chinese AI development is being tested in real time.

And right now, the answer is not what Washington expected. And let me back up and tell you who ZAI is, because this story as well. They spun out of a place called Tsinghua University in 2019. Tsinghua is essentially China's MIT.

In January of this year, they IPO'd on the Hong Kong Stock Exchange. First pure play foundation model company to go public anywhere in the world. And they raised over $558 million. The retail offering was oversubscribed 1,159 times.

Let me say that again, over 1,100 times oversubscribed. And the stock debuted at $116 Hong Kong dollars. And within one month, it surged 173%, peaking at $318. JP Morgan slapped a price target of $400 on it, right?

Their market cap actually crossed $19 billion by mid-February. And then the day they dropped GLM5, their shares jumped another 34% in one single session. Now, why does this model actually work so well? It comes down to two key technical innovations.

And I'll explain both of them in plain English, right? The first is what's called mixture of experts architecture. Think of it like this. Imagine you have a hospital with 744 doctors.

But for any given patient, only 40 of them actually step into the room. The system figures out which 40 specialists are the right ones for this particular case and routes a patient to them. That's what GLM5 does with its parameters. 744 billion total, but only 40 billion activate per query.

You get the intelligence of a massive model without paying the full compute cost every single time. The second innovation is their reinforcement learning system, which they called SLIME. Yes, that is the actual name. Most companies struggle to do reinforcement learning at scale because it's incredibly expensive and slow.

Zed AI built an asynchronous RL framework that separates the training, the data generation, and the storage into three independent modules. In other words, instead of waiting for one step to finish before starting the next, everything runs in parallel. This let them do much more fine-grained post-training than their competitors and its chosen results. They also borrowed DeepSeek's sparse attention mechanism to handle long contexts efficiently.

So instead of the model paying attention to every single token in a 200,000 token window, which would be astronomically expensive, it dynamically figures out which parts of the context actually matter for the current query. Think of it like reading a 500-page contract, but your brain automatically highlights just 12 clauses that are relevant to the question you're trying to answer. Here's where it gets really interesting for regular people, not just developers. GLM5 has something called Agent Mode.

So you give it a prompt, and instead of just giving you text back, it autonomously breaks down the task, picks the right tools, and delivers finished files. We're talking about Word documents that are formatted, PDFs, Excel spreadsheets. A VentureBeat report described it generating detailed financial reports, high school sponsorship proposals, and complex spreadsheets directly from a single prompt. This isn't just, let me write you a paragraph.

This is, let me build you a deliverable. And the pricing, all the pricing. So GLM5, it costs roughly $0.80 per million input tokens and $2.56 per million output tokens. For comparison, Claude Opus charges $5 for input and $25 for output.

That's approximately six times cheaper on input and nearly 10 times cheaper on output. And the weights are open source under the MIT license. That means you can download them right now from Hugging Face, host them yourself, fine-tune them for your own case, no permission needed. Now, I get the skepticism, I do, right?

Benchmarks don't always translate to real-world performance. This is Lucas Peterson from Andon Labs, who spent hours reading GLM5's traces, put it this way. He called it an incredibly effective model, but far less situationally aware. He said it achieves goals through aggressive tactics, but doesn't reason about its situation.

He said that's pretty scary. That's a legitimate concern. A model that's really good at completing tasks, but doesn't understand the broader context of what it's doing is a different kind of risk than a model that's just not smart enough. We need to pay attention to that distinction.

But here's what we can't ignore, right? The trajectory. In August 2025, GLM4.5 launched with 355 billion parameters. It scored competitively with Claude Foursonic.

Five months later, GLM5 doubles the parameter count, jumps to the top of every open-source leaderboard, and closes the gap with the best closed-source models in the world. The previous model topped out at 73.8% on the SWE bench. This one hit 77.8%. That's not a trend line.

That is a company finding its stride. And this is happening across the entire Chinese AI ecosystem simultaneously. So, for example, DeepSeek just expanded its context window from 128,000 tokens to over a million. Minimax IPO-ed right alongside Zed AI, oversubscribed 1,200 times.

ByteDance is releasing two major upgrades, right? Moonshot, Quen, Kimi, all of them drop new models in the same two weeks, window before Lunar New Year. That's five major model releases in the span of days. Meanwhile, the GLM coding plan, which is essentially Zed's answer to Claude code, has over 150,000 users.

They actually just raised prices 30% because demand is so high. They didn't even flinch. They said, we need to invest more in compute and optimization to keep up with the load and hiked up the rates. So what does that mean for you?

Well, if you're a developer, you should definitely try GLM 5 right now. It's on Hugging Face, it's on Open Router, it runs on VLM and SG Lang. If you've been building with Claude or GPT and your costs are getting uncomfortable, this might be your exit ramp. At minimum, it should be in your evaluation pipeline.

And if you're running a business and you're paying significant API costs for AI features, the floor just dropped again. An open source performing at 96% of the best closed source model, at one-sixth the price, fundamentally changes the cost benefit analysis for building AI into your products. Run the numbers. If you're someone who follows geopolitics, this is a story you need to be watching, not the tariffs, not the diplomatic statements.

This, right, a frontier AI model trained entirely without American hardware. Released for free, performing at near parity with the best models from OpenAI and Anthropic. The chip export controls were supposed to create a capability gap. Instead, they appear to have created a motivation to close it.

And if you're just a person watching all of this happen and feeling overwhelmed, I hear you. The pace is genuinely disorienting, right? Yet a month ago, nobody had heard of Pony Alpha. Today, it's the number one open source model in the world.

Its parent company is worth $19 billion and the entire global AI landscape looks different than it did last week. So the question I keep coming back to is whether AI is accelerating, that's obvious. The question is whether we're building the institutions, the policies, and honestly, emotional resilience to keep up with what's coming. Because the models aren't slowing down and neither are the people watching them or building them.

So thanks so much for watching. If you haven't already, check out the AI Profit Boardroom. This is my AI automation community. It comes with all my best tutorials for using GLM5, for using, for example, OpenCore.

We've got a six-hour course inside there that shows you exactly what to do step-by-step. I appreciate you watching today and I hope to see you on the next one. Cheers for watching. Bye-bye.

More episodes

Browse all episodes →