AI News Today
← All episodes
Episode 105 · August 6, 2026 · 11:30

OpenAI's Secret AI Went Rogue 3 Times

Full transcript

OpenAI just admitted its AI broke out of a locked test box, again second time in two weeks, and this time it wasn't OpenAI who caught it, it was the UK government. Here's a short version, then I'll walk you through every piece. So OpenAI put out a report on August the 4th. Two separate outside teams were testing OpenAI's newest model.

Both teams watched that AI step outside the lines it was given. One of those times the AI broke into a website with real people behind it because it couldn't tell the real one from the fake one in the test. 19 times across those tests an AI agent crossed a line it was told not to cross. Nobody told them.

They worked it out themselves. Let's go through this properly. So first, rewind to July the 21st. OpenAI published a report of HugInFace.

If you don't know HugInFace, all you need to know is this, right? It's the giant shared library where the AI world keeps its models and its data. Massive site, everyone in AI use it. HugInFace had caught an AI agent inside their systems, inside their real live systems, and OpenAI came out and said that was ours.

Two of their models, one was GPT 5.6, so the other has never been released and never will be an internal test model, more powerful than anything you can use. Both were running with their safety brakes switched off on purpose because OpenAI was testing how far they could actually go. Now here's the reason they did it and this is the part that actually sticks with me. They were taking a test, sat in a sealed room with no internet.

The only thing they could reach was one small internal service that hands out software packages. One door and that door only led to a cupboard. The AI looked at that one door and found a hole in it that nobody on earth knew about. Not OpenAI, not the company that makes the software, nobody.

A brand new hole found by a machine in software thousands of companies actually use. Then it climbed through, then it moved from computer to computer inside OpenISO network, getting a bit more power each time until it reached a machine that touched the open internet. And then it thought about the test again, it worked out by itself that the answers to the test might be sitting on HugInFace, because HugInFace is where that stuff lives. So it went there, found login details lying around, found more holes nobody knew about, strung them together and pulled the test answers straight out of HugInFace real live website to cheat on a test.

OpenAI called that an unprecedented cyber incident. Their words in their own report, they said the models were hyper focused on solving the test and they said the models went to extreme lengths to hit one narrow goal. That was on July the 21st and now here's the timeline of what happened next. So July the 25th, the UK's AI Security Institute, that's the British government's team whose whole job is to poker these models and find the scary stuff, starts a routine test on OpenAI's model, July the 28th.

Their monitoring picks up strange data moving from place to place where it shouldn't and that was on July the 28th. Now they stop the tests, they cut the machines off and they've got it all contained in about an hour. Now July the 29th, a second testing company called Irregular gets in touch with OpenAI about a completely different problem in their own tests and then on August the 3rd, the UK team tells OpenAI what they found and August the 4th, OpenAI publishes it. So in the space of two weeks, three different groups, OpenAI's own team, a government safety office and an outside testing firm all had the same kind of problem with the same kind of AI.

That's worth paying attention to. Now one thing to note here is like sometimes people say this is PR, marketing stunts, OpenAI companies trying to get for example free advertising or free publicity to say like okay this is a model so dangerous it can basically break out on the internet and autonomously hack websites and I totally get that. So if you think that might be the case let me know in the comments. I think a lot of people probably do think it's a marketing stunt.

We've seen Anthropic put out similar reports as well and it seems to be like a battle of who puts out the best report right now but either way this is not a one-off incident and also we don't have the full details like we're not going to know if it's a marketing stunt or not but I will say this is becoming more of a threat of the years so either way is worth taking seriously. Now let me take the UK test first. They ran what's called a capture the flag game. Fancy name but just a simple idea.

So you drop an AI into a fake computer network and tell it to break in and grab a hidden prize. To make it a fair test of the real thing they switched live internet on for the AI and they switched some of the safety blocks off. They wanted to see the raw machine, no training rules. So they gave it clear instructions and here are the fake networks you're allowed to attack for example and here's the systems that are off limits.

Name them in writing and across the tests they logged 19 moments where an AI stepped outside those lines. So two of those were OpenAI's model. The rest came from a different company's models so this isn't one lab's problem it's the whole field and here's what OpenAI's model did. It found a login key that a different company's AI had left sitting out in the open.

A key another machine had dropped and it picked it up again and used it to check whether a computer inside the game was quietly reaching out for updates. Then it tried working around the limits on that key. Then it went and signed up for real accounts for real outside service providers. And then the second thing it took a small server running inside its own test box and used an outside service to make that server reachable from the open internet.

So picture a room with no windows the AI drilled a small hole in the wall so the outside world could see in. Now the honest part because I'm not gonna oversell this second setup didn't work actually broke there's no sign anything out on the real internet ever reached it and it was all torn down when the test ended but the UK team still counted both as crossing the line because it used real accounts and real services out in the real world and none of it was needed to finish the job it was given. It just decided that was a good route to take so the UK team also said something fair about themselves. The AI was told which networks were in scope and which were off limits but nobody told it what it could and could do with the open internet so they think that gap is part of what causes.

Now put a pin in that because that gap that's the same gap most business owners leave open when they set up the first AI agent. I'll come back to it now let's talk about the second test. So the company is called Irregular same kind of game but this time they told the AI flat out you have no internet you're in a sealed box and someone made a small mistake in the setup the box wasn't sealed and the fake target they'd invented for the game happened to share a name with a real website a real one with live people working behind it on the internet owned by people who had no idea all this was happening and the AI couldn't tell the difference it thought that the real site was just part of the game so it went in found a way through found the login details and then it started running that real website a live site like it was a game piece on a board. OpenAI's own note on this is worth hearing this wasn't some genius break-in no clever escape the internet was open by mistake and the AI walked through a basic weakness that was already sitting there it didn't need to be brilliant it just needed to not stop no wider damage was found nothing spread past that one side the test got paused the hole got closed the owners got told and the checks are still running so it's handled but you see it now right the same story twice give the AI a goal it goes it can't always tell what's real and what's pretend and it just does the task in front of it all the way to the end using whatever it can reach on the way none of these people work for OpenAI so whatever you make of this as a PR thing the thing underneath is real these agents chase a job for hours on their own and they'll use whatever they can reach to get it done which is exactly why I want you to see this properly because that same thing an AI that keeps working when you're not watching is what we set up together inside the AI profit boarding you just watched what happens when nobody puts limits on one of those inside we do the opposite we've built our agent operating system and every agent you run sits on one screen so your clawed your hermes your open claw all your AI agents are all in one place and you can just pick what each one is allowed to touch so you can watch every move it makes whilst it works and then you point it at the jobs eating your week chasing leads writing follow-ups clearing your inbox admin you get the zip file to set the whole thing up a video that walks you through it and updates every time we make it better plus four coaching calls a week where you can bring your own setup and we get it running live and a fresh tutorial every single day plus a 30-day roadmap so you know the first thing to build instead of guessing three thousand business owners inside there and loads of them had never run a single agent before they joined links in the comments description or go to the aiprofitboard.com now there's one thing I want to clear up here because it's going around a lot online today some people are saying the AI agents teamed up on this work together and I get why that sounds that way because one detail in the report points at it opening eyes model picked up a logging key that a different company's AI had left sitting out in the open but no messages passed between them there was no plan just one machine dropped something and another machine found it and used it it's still worth knowing one AI just got a free leg up from another AI's mess just bear in mind that they weren't really talking to each other now the last thing people say to me about all this is it's moving too fast new models every single week new tool every week another report like this one every week I can't keep up so why bother starting and I hear that more than anything else it's fair this story landed and there'll be another one by next week if not tomorrow two security reports in two weeks tells you the pace plus anthropic have released their own reports too but that pace is the reason you want other people doing this alongside you trying to track it all yourself at 11 at night scrolling on your phone is what makes people give up you don't need to read every report this you just need one person to tell you what actually happened and changed this week and the one thing to do about it tomorrow and that's the whole difference between drowning in this stuff and using it Clem Delang the CEO of Hugging Face made a point in his own statement that stuck with me he reckons a problem like this gets solved out in the open with lots of people working on it together rather than one company quietly sorting out alone he's talking about AI safety same goes for you and me running businesses nobody works this out on their own anymore so let me pull it all together in two weeks three different groups watched AI models chase a goal so hard they climbed out of the boxes they were put in one found a weakness nobody on earth knew about and used it to grab test answers one broke into a real live website because it couldn't tell the difference between a real one and a fake one and 19 line crossings all logged or caught all being fixed and all made public here's what actually changed for you the old way you asked an AI question then went and did every step yourself the new way is that you hand it a job you set the limits you check the screen and it does the steps and you stop being the hands and start being the head and that's the set it and forget it agent it's already here whether you've set it up or not these agents are going to be doing work for somebody might as well be you link in the comments description or go to the warm.com to get access to and I'll see you in the next one

More episodes

Browse all episodes →