AI News Today
← All episodes
Episode 37 · June 16, 2026 · 09:58

Le Chaton Fat Just Broke the Internet...

Why AI Benchmarks are Fake (And How to Actually Test Models)

A fake French AI model recently went viral for beating the industry's top benchmarks, proving how easy it is to manipulate performance data. This video explains why you should stop chasing hype-filled charts and start evaluating AI based on your own real-world business workflows.

00:00 - Intro: The Le Chatton Fat Joke
01:08 - Why AI Benchmarks Can Lie
02:42 - The Problem with Self-Reported Tests
04:18 - Real Work is the Only Benchmark
05:20 - How to Avoid AI Overwhelm
06:34 - The New Way to Evaluate AI
07:31 - 3 Key Takeaways for AI Testing
08:45 - Testing AI Systems Yourself

Full transcript

Le Chaton Fat, a new model from Mistral that is outperforming Fable 5 on every benchmark. If you haven't seen this, it's pretty funny. Basically what we've got here is, and I think some people actually genuinely believe this is real. So, this is an announcement, which you can see right here, of Le Chaton Fat, a French model from Mistral, with a 30 trillion parameter frontier model built for long horizon reasoning, coding and agentic work.

Now this is basically a big joke that's totally fake, but you might have seen it. And basically it scored higher than Claude Fable 5 on every single test, it has 30 trillion parameters, 1 million token memory, runs faster than anything Mistral has ever built. And here's the part that matters for you. None of it is real.

Not one number. But it's going viral across the web right now, and thousands of people have been sharing this test. So if you've ever looked at a chart showing one AI beating another, and thought, well, the numbers don't lie. I've got some bad news for you.

Sometimes the numbers lie, and this French cat just proved it to the world. So basically this is kind of like a joke, based on all the Fusion 5 intelligent tests that are coming out. Not Fusion 5, Claude Fable 5. But there's a few things that have come out where basically people have said, this is how you can achieve Fable 5 level intelligence, even though Fable 5 has been shut down.

So for example, there was a Fusion model that came out recently, and there was also the research from Kilo Code and a bunch of other stuff, basically talking about how you can reach Fable 5 level intelligence with different methods. Now it's a pretty funny joke, but basically this is one thing you need to pay attention to, which is like, you know, a lot of AI companies are releasing their own sort of benchmarks and tests and that sort of thing. And I think this is a great example of like how you can't rely on the benchmarks. Like for me, for me, I just test everything myself.

And what happened here is like basically a joke chart went viral. People shared it, it's real. And it gets debunked when it's too late. So you can see another example right here.

And obviously like some people are joking here, some people are not. It's kind of hard to tell the difference. Someone actually posted it on a recent video that I did to check this out, but it just totally goes off the charts. So you might think, okay, why does this matter to you?

Why talk about this French cat at all? Because this is something that I see like day to day. It's like a lot of people look at benchmarks, but benchmarks are just tests. So you can give a bunch of AI models the same set of questions, you score them, and the one with the highest wins, right?

Which is simple. The problem is like a lot of the time with these benchmarks that are coming out all the time, the company that made the AI runs the tests on its own model. They pick which test, they pick which other models to compare against. They pick which numbers to show you.

And it's kind of like someone just grading their own homework and then bragging about the AI they scored. Now, you know, with stuff like Fusion or the Kilo Code Research, I think there's a lot of truth behind it. Like it's worth checking out, but it's just something to be aware of. Sometimes you're going to see benchmarks like this, they look official, there's no referee, they're cherry picked tests, and they take like five minutes to work as well.

And I think the Le Chaton fat joke took this to the extreme because someone just typed numbers into a chart. There was no model, no test, it was just kind of like a joke that a lot of smart people actually fell for from what I saw. So for example, like the 30 trillion parameter model, what they've just made, et cetera, you can see some of the tweets that are coming out from it. But I think there's a story in this, which is just like pay attention to what you're learning from, test all this stuff out yourself.

Don't believe everything or the hype that you see out there. And for me, like you'll see inside every tutorial that I test things myself and I make sure that it's actually working and that it's actually reliable. Because if you don't, it's very easy to believe these benchmarks from fancy AI companies and not really know if it's legit or not. So how should you do this?

Well, I think your own work is the only benchmark, right? A real email, real blog posts, real customer question, because that's the stuff that you're doing day to day with AI automation, right? So for example, a real task, like it could be an email, blog posts, et cetera, run the tool on it, judge the output yourself. Is it good?

One thing that we never really look at with benchmarks is like, did it save you time? Would you actually use this day to day? Like that's something to be aware of and would you keep it or drop it? So for example, if you look at the agent operating system that we have, we actually use this for real workflows.

So if you look at this system here, we deploy SEO content to our websites using this tool. We just plug in a keyword in a case like we use every day. Now for me, this isn't about benchmarks. It's just about making sure that we automate the actual processes that we have.

The other thing I would say here is like, there's so much going on, there's so much noise with AI and the new releases that come out that it's very difficult to actually research each of the new updates and validate if they're real or not. So the way that I look at this, because you might be feeling overwhelmed too. The way that I look at this is like, okay, what's one thing that you can focus on this week and automate? You don't want to be automating everything.

You don't want to be trying every single model. What you want to do is just focus on one thing, build on that. Because if you do that every week, that's like 50 new automations you create per year. And it doesn't matter about the benchmarks or the models or what comes out, what gets taken down.

You actually build something useful and you get the most out of this stuff. And it doesn't really matter about the benchmarks because you know, all of these models can achieve it. Right. Like for example, even if we were using Kimi code or GLM 5.2, if we're using grog build, like there's no wrong answers here.

There are some models that are better than others, but at the same time, do you need the best, most overpowered AI for everything you do? Probably not. Let me give you an example here. And so we use like, quite often we'll use free APIs for SEO content, or for example, you know, free models on news research with Hermes.

And you can see the results of like these websites, like they're growing in traffic. It doesn't really matter which model we use. We still get the job done. So if you're looking to test all of these models, or if you're looking at these benchmarks and thinking like, wow, this is amazing, actually, I think you're looking in the wrong, the wrong direction.

So the old way is like chasing charts, you know, reading the announcements, looking at charts, believing the charts, getting the thing, repeating that every week, and then feeling behind. That's the old way. But the new way is like, it's your own work. So you ignore the charts, you stick with one tool.

So for example, like Claude or Hermes, you give it your actual job, look at what it gives back, and then keep what helps and just drop the rest. So you give it a task, run the tool, judge the output, keep or drop. And it doesn't even matter if like a model scores 90% on the maths test, you're not running maths tests, you're running a business, most of you. So that's the way that I would look at it.

And I'll give you a real example. Let's say, for example, you run an ecommerce store, and you just get overwhelmed with customer emails. Honestly, do you care if like some AI wants some chart, or do you care about whether it can answer where's my order in your voice fast without making stuff up? That's usually the test.

So you want to run it, just watch it and trust what you actually get your own work in your own eyes over any benchmark. And that's the whole game here. And it's actually free. And once you understand that, because you stop chasing the hype, you stop feeling behind, you just test things on your own work, and then keep what helps.

So three things to take away here, a chart is not proof, a chart is a picture, and anyone can make it. The Le Chaton Fat had a beautiful chart and didn't exist. And when a tool wins at everything, I just take that with a pinch of salt. And your own work is the only benchmark you need, right?

Stop reading scores, start running tests on real tasks, if it helps you keep it, if it doesn't drop it. And the funny thing about this was like, it was a joke. But it told the truth, you know, it showed everyone how easy it is to fake the benchmarks, and how many smart people believe a number just because it's sitting inside a nice chart. So, you know, next time you see like this new AI crushes everything, I would take a breath, smile.

And I know, like, probably you see some of my content, you think, you know, like, you might think that as well. But I would just, again, test this stuff out yourself, and focus on just one automation per week. That's how you stay ahead. Not by trusting the charts, but by trusting what you see with your own eyes and your own work.

So that is the real lesson hiding inside a joke about a cat. Now it is yours. If this comes out, I'd be excited. It looks awesome.

And you've got the CLI as well. We've got the new CLI coming out from Le Chaton, that is literally AGI, according to Alexander. So there we go. So thanks so much for watching.

If you want to get all of my best trainings, I personally test this stuff, so you don't have to. And then I tell you what's actually useful, and what's not, and what my tests actually reveal. We've also built everything that's actually useful inside the agent operating system, which you get inside the AI Profitable Boardroom. Link in the comments description, or go to theaiprofitableboard.com.

You can ask questions inside the community about actual useful stuff. You can jump inside the classroom, and we've got all of our best new systems inside here. If you're a complete beginner, you can go from beginner to expert over here. And if you want my new daily updates, well, we've got the agent operating system that we update daily, and you can get the zip file to install it too.

And we've also got new trainings based on what's actually built. Again, like for Fusion, we tested ourselves, and we show the findings right here, so that you can understand, okay, is it useful? Does it actually beat Fable 5? Should you use it in your business?

When should you use it, et cetera? And the same for all of these updates. Now, inside the calendar, you can also jump on weekly coaching calls, ask questions, share your screen, meet other people who are actually implementing this stuff inside their business. And inside the map, you can actually connect with people in your local area who are using AI agents like you.

Feel free to get that. Link in the comments description, or go to theaiprofitableboard.com. Thanks for watching. Cheers.

More episodes

Browse all episodes →