Full transcript
Inkling, it's a brand new model that's just dropped. This is a four weights model, open weights. Um, I'm going to be testing out today. I'll show you some of the stuff that we've tried with it.
So you can judge for yourself how it performs. Now, this is a really popular model. It's just blown up. So it's called Inkling.
It's a mixture of experts, 975 billion total parameters, 41 billion. Supports a context window of 1 million tokens, pre-trained on 45 trillion tokens. And it is the first in a family of models of different sizes, which I'll run you through today. And it's also available directly inside open code.
So you can actually run Inkling inside open code. That's actually how we tested it. And I'll show you what we built in a second, how it performs. Now, one of the most important things to note here is it's available for fine tuning on Pinker, and you can start using it inside the Inkling playground.
So how does it perform? Well, number one is available for free on Hugging Face as well, as you can see. So it's an open weight project on Hugging Face. And I've called this framework, the Ink Machine.
So there we go. And let me get straight into this. So if we go into open code inside our Egentic OS, you can see that we can switch between the different models. We're going to start with Inkling and then we're going to try this out.
So we'll say, for example, create a beautiful landing page for an SEO agency, just to show you what it looks like when it comes to creating and designing web pages. Really, this is not designed to be like frontier level. You know, I don't think this is going to compete with Fable 5 or anything like that. But at the same time, it's open source.
If you want to see what we've built with it so far, you can see some examples. So these are all different examples we've created with Inkling. The font is not too bad at all, actually. So this is our workspace.
These are all built with different prompts. We open this up. Very visual model, like it creates some beautiful stuff, as you can see here. Let's open this one, kind of like a, it moves wherever I move my mouse.
It looks beautiful. Here's another example of like a kind of neon clock that we have created. What we got over here. This one was not particularly good.
This was not good. Just going to be a hundred percent honest with you. And then we have the matrix wallpaper style theme here. That was pretty nice.
The Inkling fireworks just crushing it. Absolutely crushing it. That looks insane. And then we've got the Inkling orbit style as well.
Now you might be wondering, okay, how does this compare against all the other models out there? So we've got GoldieBench. We test about 50 projects with everything that we try. And this is GoldieBench.
So you can see, for example, like Fusion, Hermes Mixture of Agents, GPC 5.6. So they're right at the top of the leaderboard. Now, Inkling is coming at 14 in total out of about 30 different models. So it's scoring six on average out of 10, which means kind of mediocre from what we've tested it so far.
Actually, it's a new model from Thinking Machines, but it is outperforming stuff like, for example, Nemetron 3 Ultra. And so it's not a bad open source model. I wouldn't say it's on the same par as, say, something like GLM 512, which genuinely felt like Opus 4.8 level. But, you know, the benefits of this are that it's open source.
You can create some nice one-shot stuff. It's actually scoring pretty high on Frontier-class agentic coding for an open model, and it's a million token context window. So those are the good things. What are the bad things?
Well, actually, and you'll see in a second, some of the stuff we've built with it, particularly when it comes to like 3D games and stuff like that, pretty bad. It's not great at like physics and that sort of thing. And it's certainly not the strongest overall when it comes to, you know, Frontier models or open Frontier models. So I've shown you the good stuff.
Let me show you the bad stuff now. Here's a few examples of what we tried to create with 3D games. And these are just standard tests we run for every model on GoldieBench. So this was a Crypt Dungeon game and it didn't work at all.
This was like a, it's called Dragon Realm, as you can see right here. And we can try and move it around and that sort of thing. But it's not really that interesting at all. And then this actually looks decent when you have it like that.
But if we try and move, what's going to happen over here, mate? You know, so there was a lot of stuff that totally broke. I mean, we tried to create this Skyrim style game here as well. It's just not good at like 3D games, basically.
Let's have a look how it performed on that website that we just built a second ago. So we've created that over here. Let's open this up. I mean, this is, this is, this is a website, I guess you could call it that.
But it's just, it's nowhere near the same level as something like GLM 512. And I assume like that's probably where they're aiming. However, open source, I always support that. So I can't knock them for trying.
But I know a lot of people are going to ask me like, you know, should we use Inkling, should we be using Inkling instead of GLM 512? How does it perform, et cetera? It's not up there when I've tested it personally. And we've seen the outputs live today as well.
Now, having said that, apparently, according to some of the comments I read, it is the first open source 1 trillion multimodal model that supports image, audio, and text. So that's pretty impressive. But it's just not the same level as GLM 512, which is, you know, outrageously good. If you want to see how it compares against GLM 512, let me pull up the examples.
And I'm only saying that because GLM 512 is probably the best open source model that I've tried. So let's pull up the comparison head to head. So if we look at that dungeon game that we created a second ago, this is the output from GLM 512, which actually works. I mean, it's super dark, but it does actually work.
Whereas I've seen with Inkling, it totally broke. We have a look at the Raycaster style game. You can see that works beautifully with GLM 512. It's just on a totally different level compared to Inkling.
And I do think like out of everything that I've tested, China seems to be way ahead of everyone when it comes to open source models. I mean, this looks pretty cool. Whereas obviously we've got this version from Inkling, which was not great. And if we actually, this, this is like the open source.
So this is called Dragon Realm, the open world game, as you can see right here. So this is from GLM 512. Just to show you, you know, what the standard is like right now for open source models and, you know, it looks awesome. When we have a look at Inkling, you can see it's just nowhere near the same level.
Now you also might wonder, how do you actually use it? So the way that I use it is I got a tinker key. So you can get that at tinker.thinkingmachines.ai. And then you can grab an API key.
From there, you can actually wire it into open code, which we have inside that agent OS now. So you can use open code as well as, for example, any sort of model that you prefer to use with open code. And then you can start building with it directly. I do think the open source models are definitely worth checking out.
I just don't think this is the best one that I've seen. So that's basically it. That's how to use it, how it performs. We've tested it across 50 different tasks, scored six out of 10.
Not as good as GLM 512, but we do like open source. So I'm excited to see what they come out with next. I think they'll release quite a few more models. It'd be good to see if they can produce something even better.
If you actually look at the benchmarks here. So when it comes to design arena, it's actually on the benchmarks is scoring really highly. But again, this is why we created GoldieBench and it's really the only benchmark that I pay attention to because at this point you just, you can't trust these benchmarks. I don't know anyone who really pays attention to these benchmarks.
Like you just got to test everything yourself. Don't even listen to me. Don't even listen to my benchmarks. Test it yourself.
See what it's like in reality, because as you've seen today, very, very different results when you try it out yourself. And maybe that's the way I prompted it. Maybe I'm doing something wrong there. So, you know, I'm totally open to being wrong.
So thanks so much for watching. If you do want to get our agent operating system, which comes with open code plugged in. We have all sorts of amazing tools and workflows. As you can see, we have a memory system right here.
Every time a new open source project like this comes out, it's worth checking out. We actually create and add it inside our system. So you can get the most out of it. And if you want to get a full agent operating system, you get that inside the Air Profitable Boarding.
Link in the comments description or go to theairprofitableboarding.com. Inside the community, you can ask questions, get help and support in real time, and I personally answer them with a video tutorial every single day. Inside the calendar, you can drop a weekly coach calls, get help and support in real time, share your screen, et cetera. Inside the map, you can meet people in your local area who are building with AI agents like you.
And if you want to get our agent OS, it's inside the classroom over here. Go to the agent OS system over here. And every time something new and useful comes out, so for example, like Tencent dropped HY3 last week, which actually, if I compare them side by side now, I would say HY3 is probably better too. You can get all of this and all the tutorials inside the Air Profitable Boarding, so hope to see you inside there.
Thanks for watching.
More episodes