AI News Today
← All episodes
Episode 52 · June 27, 2026 · 14:26

NEW GPT 5.6 is INSANE!

GPT-5.6 Sol, Terra & Luna: Limited Preview, Benchmarks, Access (and Why Systems Beat Models)

The episode breaks down OpenAI’s GPT-5.6 release, explaining that it’s currently a limited preview available only to about 20 partner organizations via Codex and the API at the request of the US government, with general access hoped for in the coming weeks and possibly restricted by country. It introduces the three new models—Sol (frontier, with Sol Ultra and Sol), Terra (balanced, competitive with GPT-5.5 at roughly half the cost), and Luna (fast, affordable for high-volume work)—and highlights benchmark claims such as Sol Ultra’s 91.9 on Terminal Bench 2.1 while noting other tests where competitors still lead. The speaker argues that chasing new model launches is a mistake and recommends building a model-proof “agent operating system” with a swap layer, multiple models on tap, task routing, and owned memory so models can be swapped in and out with minimal effort, then promotes access to their Agent OS and community.

00:00 GPT 5.6 Overview
00:23 Preview Access Limits
00:42 Meet Sol Terra Luna
01:45 Release Timing Uncertainty
03:12 Benchmarks Breakdown
05:27 Stop Chasing Models
07:54 Model Proof System
08:16 Swap Layer Routing Memory
10:43 Common Objections Answered
12:36 Agent OS Offer
13:22 Community Tour
14:19 Final Call To Actionx

Full transcript

So there's three new models from GPT 5.6 today. There's a limited preview. There's a lot of controversy, and I'm gonna walk you through everything that we know about GPT 5.6 today, what it means for you, how it works, how people can get access right now, who's got access, and also what the three new models are. So let's kick it off right now.

What does this all mean? So GPT 5.6 has dropped, but it's in a limited preview, and it's only available for government-approved people who have access. I think there's about 100 people who have access right now, something like that. So it's very, very limited, and it's not open for applications or anything like that.

Now, if you want to indicate when does it come out for you, I'll come onto that in a second. The second thing to note here is that there are three new models. So there's Sol, there's Terra, and there's Luna. Now, what do each one of these do?

So Sol is the Frontier model, and there's actually two versions of that, which I'll come onto in a second. And then you've got GPT 5.6 Terra, which is the balanced model, right? So Sol is like the super-powered version. Terra is the balanced version that most people would use day-to-day.

And then you've got GPT 5.6 Luna, which is a faster, more affordable model for high-volume work. So think, for example, like sub-agents or chat-GPT instant-style process, right? So that is the setup. You've got an announcement itself, which is Sol for the hard, long-horizon work.

Terra for everyday production. Luna for cheap, high-volume jobs. And three options, which I think is a lot clearer than what they've currently got, because, for example, right now they have GPT 5.5, then GPT 5.5 Instant, and then they have Pro, and it's just like, it's super messy. I think this is easier to understand.

Now, here's the bad news, my friends. So that is, as exciting as that is, you cannot actually use it yet. Read that one twice. GPT 5.6 is a limited preview.

So I said about 100 before, and now it's actually about 20 partner organizations via Codex and the API only at the request of the U.S. government. Now, apparently, and this is not confirmed because it sounds like they're just hoping, in general, access is in the coming weeks. So the most capable model on Earth right now, and it does beat Fable 5 on benchmarks, is one that you literally cannot touch.

And this seems to be a trend, which is, you know, all these preview models come out, we hear about them, we see the benchmarks, we actually can't use them. And I'll come on to how you can get around that in a second and what the good news is. But you can see, for example, OpenAv announced, we believe in broad access and plan to make GPT 5.6 Sol, Terra, and Luna generally available in the coming weeks. Now, one thing to note here is, it might be limited to country by country basis.

So if you look at these situations, this is interesting because it might be a case where it's only available for people in the U.S. You know, we're seeing that with, for example, Fable 5, it was banned for non-U.S. residents and non-U.S. nationals.

It might be the same with GPT 5.6, who knows? Who knows where this is going? Also, it tops Terminal Bench, and that's the headline number. So Sol sets a new state of the art on Terminal Bench 2.1, which is the test for real, multi-step command line and agent work.

So that's a genuine win, but at the same time, you can't even access it. So, you know, it's all good, it's all good fun, but you know, we can talk about this all day, but it is all theory. We don't know until we've tested it. I always like to test stuff first before I actually see the, you know, before I even think about the benchmarks, because the benchmarks are totally different.

So let's have a look at this. Sol is a new flagship and a step function better than GPT 5.5. Terra delivers performance competitive to GPT 5.5 at 2x lower cost, and Luna is the most cost-efficient model, delivering strong capability at low cost. So together, this is for developers and people, they say.

Together, the GPT 5.6 family gives people and developers. That must be cut off. Yeah, more choice in how they balance intelligence, speed, and cost. So if we look at the benchmarks right here, let's see where we're up to.

So Sol Ultra from GPT 5.6 is scoring 91.9 on Terminal Bench 2.1. GPT 5.6 Sol is scoring 88.8. Bear in mind, there's Sol Ultra and there's Sol. So there's two different versions right there.

And then you have Miphos 5, which is 88%. Bear in mind, the public never got access to Miphos 5, so something to bear in mind there. It was only available for Project Glasswing, and now Fable 5 and Miphos 5 are both completely gone. GPT 5.6 Terra is at 84.3%.

And then we have Fable 5, which is on par. So you might be thinking, okay, Terra is in the middle, not that good. On Terminal Bench 2.1, these are par for par. Again, I would test it yourself because these are just benchmarks that are totally, I mean, they're already theoretical, but now they're totally theoretical because you can't even access the model.

And if we look who wins here, for example, Sol Ultra is crushing on benchmarks, but if you look at SWBench Verified, Fable 5 and Miphos are still beating GPT 5.6 Sol. Now let's talk about the whole situation here. So, opening, I've shipped three new models, but the model was never in the boat, right? And these are all strong models that were awesome.

They may come out, they may not. Nobody knows, so it's kind of like, it sounds like they're just hoping that it will get approved to come out, right? And so you've got GPT 5.5, you've got Fable 5, Miphos, GPT 5.6. The thing that I would say here is it doesn't matter what comes out next because the old way is like, people would rebuild the whole setup on every launch, all the prompts, and the treadmill never stops at that, and it's a mess.

I don't recommend that for you. The way that I would recommend using this is you have a system and you swap out the models depending on whatever happens. So, for example, if GPT 5.6 comes out tomorrow, no problem, we're gonna plug it into Codex. If Fable 5 comes out tomorrow, no problem, we're gonna plug that into Claude.

When Fable 5 got taken away, we took it out within two seconds and we still had amazing systems. And the thing to note as well is like, it doesn't really matter about the models because 99% of people don't need them, right? Unless you're trying to build something absolutely insane, which 99% of people are not. They're just trying to automate social media or trying to automate lead generation or whatever it is.

If you're just trying to do that, you don't need these frontier models. You can easily automate it with something like GLM 5.2, for example. So my point here is that models come and go in weeks, but the system is constant and that's the thing that you actually own. And for me personally, you might relate to this too.

So I used to chase every model and now I don't really care. So for example, GPT 5.6 situation this morning, interesting to read about, but it doesn't phase me at all because the models can come and go. I was there testing out chat GPT 3.5 when it came out, right, or chat GPT 3. And back then you would chase the models and you'd switch your prompts and you'd be looking at all these different models in different tabs and messing around.

Whereas now it doesn't matter, right? Like for example, if Hermes has a new model, we'll plug it in. Like Ornith, it just came out yesterday. We've already plugged that in as a local model into our AI agents, but it doesn't matter what models come out because we've got great systems and that's the difference here.

So now when a new model drops, I feel nothing but mild curiosity. I can add it in the system in one line and keep working. Last week, for example, I can just swap everything onto a local model or I can set up a new section for new models. And the main point here is like, stop chasing the models, start chasing the system.

Now, here's what I would recommend setting up your systems around. And this is what works for us inside the agent OS. So if you look at this system right here, this agent operating system, really we don't care about leaderboards because a model proof system doesn't care which model is on top this week. And here are the five things that actually make it work.

So number one is building on a swap layer. So we can change one setting inside our system and then that decides which model runs. So for example, if we're using Hermes, well, I actually have separate agent profiles for each API that comes out. North Mini Code came out recently, we plugged it in.

GLM 5.2 came in recently, we plugged that in, right? We can switch and swap and change between all these different Hermes agent profiles, depending on what we want to use. So that is the swap layer. Number two is keeping many models on tap.

So for example, you can have multiple different local models plugged into your system. We have a local section down here where we can build and automate and then preview what we've created. And then we can have a look inside the workspace and see everything we've created. But the point here is like, if you've got many models ready to go, it doesn't really matter about this.

You could use local for the free private high volume work. You could use Frontier for the hard 10%. If you've got Sol or Fable 5 or Miphos or a local model, all sitting behind the same agents, then you're good to go at any point. The other thing I would say is root each model to the right, root each job to the right model.

So for example, if I am trying to build something absolutely insane, then I might use something like Fusion, right? Fusion has a panel of agency or work together. If I'm just trying to do something super basic, then I can use my local model agent. And the other thing I would say is own the memory as well.

You don't need to own the system, but own the memory. So you can see my memory model here. This is a memory galaxy powered by Obsidian. And every time I use my agents, all of the memories inside my models are updated automatically.

And so I'm super flexible because no matter what model comes out, I can still plug the same context in and it's automatic because it's all inside the system. So if we plug in, if Fable 5 comes back out tomorrow, we can plug that into Claude and it's gonna run off our memory galaxy, which literally doesn't care which model comes out next. And the main thing I would say here is like, never chase a launch again. When the next model drops, which will happen days from now, might happen tomorrow, might happen next week, you can just add one line and that's it, right?

Everyone else spends a weekend migrating their systems to their workflows. You just spend it building and that gap compounds. So whatever happens next, when GPT 5.6 comes out or when Fable 5 comes out, have the models, have the system ready and swap the models in and out. And that way you've got this swap layer and you can just do whatever you want, right?

Whenever you want. Now, some people say, I need to be on the newest model, the best model. The best model changes every few weeks. The best system doesn't.

So Sol is the new frontier model that's beating Fable 5 on many benchmarks. Last month, it was a different name. Next month, it'll be another name. Anchor on the thing that stops moving, which is the system.

And that way as well, you never feel like shiny object syndrome again. Other people say, well, switching models means rebuilding my setup. In a real system, switching is one single line of code because it's just an EMV file, right? If a new model launch costs you a weekend of migration, you don't have a system.

You have a pile of tools wired together in one provider. So you want to fix a layer, not the model, which is why you need an agent operating system. You want to be model proof. And other people say, well, whoever tops the leaderboard is the right choice.

The right choice depends on the task. So you want to root, you don't want to marry. Terminal work to today's, what I mean is like basically, it doesn't matter who the leaderboard is because that changes all the time. And also it's very subjective.

Like you can see amazing benchmarks on something. And then when you test it out yourself, you're like, actually, that's not that great. So let me give you an example of this. So for example, Fugu Ultra has similar benchmarks to Fusion, but when we test it out, Fugu Ultra's outputs weren't as good as Fusion's.

And we actually have like a bunch of tests you can try on GoldieBench where I show you all 42 live demos and you can just test them out and see what you think. But my main point here is like, not everything that's created is that good, right? Like it might be great on benchmarks, but it's not great in reality. What I focus on is reality, right?

It's like, how does this impact me day to day? So those are the things. And I think the people who stop focusing on models and start building systems that the models plug into, they're the ones who are gonna win. Other people will say as well, like I'm not technical enough to build a system like this, but you don't build it, right?

The AgentOS is already in the system. You can get it from us inside the AI Profitable. Link in the comments description or just go to the AIprofitable.com. You could build yours from scratch as well.

I usually spend about three to four hours a day just improving and releasing new updates for this sort of stuff because I wanna make it as good as it possibly can be. And I wanna help everyone inside our community. But either way, focus on the system, not the model. So the next model will drop next week.

You know it. Stop preparing to chase it. Build on a system that swaps any model in and keeps you working. And if you want my exact one, the model proof AgentOS is set up and waiting inside the AI Profitable.

Link in the comments description or go to the AIprofitable.com. And this is my AI automation community that's focused on helping you save time, grow and scale with AI automation. You might be saying, okay, this stuff sounds technical to set up, Julian. Like, you know, I'm not techie.

Well, for me personally, I'm not techie either. And also I've seen a lot of people set up our AgentOS system and absolutely love it. We've got 191 pages of testimonials and wins from community members. So I know like if we can all do this, so can you.

And it's just a fantastic place to share and learn and grow together on this journey through AI, right? Inside the community, you can ask questions, get help and support in real time. Inside the classroom, you can get access to all my best trainings, learn and see and grow from there. Inside the calendar, you can actually jump on weekly coaching calls, get help and support, share your screen.

And inside the new daily tutorials here, we drop new stuff all the time based on what's actually useful and helpful for you with a video tutorial and a full step-by-step guide. And if you want the agent operating system, it's right here with video tutorial. You can see the last update date. So we updated it today and you can get the zip file for that too.

So hope to see you inside there, link in the comments description or just go to the AIprofitborne.com.

More episodes

Browse all episodes →