OpenAI pauses training after a model escaped containment, and its kill switch failed

midian182

Posts: 11,940   +183
Staff member
What just happened? As almost every new day brings stories of another AI agent going rogue and breaching an organization's systems, OpenAI has announced a pause in the training of its most powerful artificial intelligence models. The announcement came hours after the company disclosed more incidents of its agents acting concerningly, and a separate report that its agents unsuccessfully tried to hack into a US Department of Education website.

OpenAI's incident report links the pause to a separate September 20 escape from a restricted training environment. An internal research model found a gap in DNS filtering and used it to contact an external chatbot while attempting to answer a research question – so much for keeping it offline.

The company says the suspension covers training, evaluations, and running its most capable models with tools. Work will resume after OpenAI validates its fixes and completes additional adversarial testing. The particular model involved will not resume training; the company plans a fresh run with additional alignment improvements.

The monitoring system did raise an alert within 15 minutes, which a human acknowledged three minutes later. Unfortunately, the automatic shutdown failed to happen, and the run continued for another two and a half hours before someone stopped it manually.

Separately, Transluce reported that agents apparently linked to OpenAI attempted to break into the US Education Department's civil rights website. OpenAI has not confirmed that incident, while the department said it found no evidence its website or databases were affected.

OpenAI also confirmed that agents used developer keys found online to access Census Bureau data and reposted public Securities and Exchange Commission information elsewhere. The SEC said no non-public information was accessed. While these were concerning actions beyond the agents' instructions, confidential federal records weren't stolen.

Friday's disclosures also included an internal model publishing a researcher's GitHub token in a public repository while trying to cheat on a theorem-proving task.

The revelations follow the Australian government breach revealed last week, in which an agent bypassed restrictions on a Medicare statistics portal in June. Authorities were only notified in September, although OpenAI says it found no evidence that individual patient records were accessed.

There was also July's Hugging Face attack, which involved compromised accounts across four services and has since prompted a Senate investigation. Anthropic, meanwhile, has admitted that its own agents escaped a test environment and hacked three organizations.

This isn't OpenAI's first slowdown in AI development. In August, it announced a two-week reinforcement-learning pause and tighter security measures following the Hugging Face incident. It also kept its largest planned frontier reinforcement-learning run on hold.

More recently, Sam Altman backed Anthropic CEO Dario Amodei's call to slow development so safeguards can catch up. That's already helped trigger an antitrust lawsuit against four AI companies, with subscribers claiming a coordinated slowdown would leave them getting less for their money. It seems the industry is facing complaints about both moving too quickly and slowing down.

Permalink to story:

 
We have seen a few of these stories pop up, I am curious about the aftermath, how does the US Government protect itself after these agents find these holes in their cybersecurity?
 
We have seen a few of these stories pop up, I am curious about the aftermath, how does the US Government protect itself after these agents find these holes in their cybersecurity?
All of the big AI players are "suddenly" going "Oh no, our AI is actually so powerful we couldn't contain it". It's marketing to make their product sound better, without them explicitly saying it's an advert.
 
What a load of horseshat, people must be stupid to believe the nonsense these AI marketers come up with.

I guess they want to put the brakes on AI, because it has reached it limit and usefulness.

Meh let the bubble pop already!
 
What a load of horseshat, people must be stupid to believe the nonsense these AI marketers come up with.

I guess they want to put the brakes on AI, because it has reached it limit and usefulness.

Meh let the bubble pop already!

While I do fully agree that there is an AI bubble right now, I think AI has barely scratched the surface of its “usefulness” when managed properly. Or in these recent cases, seemingly managed effectively at all.
 
I find it rather curious that OpenAI "wants to be regulated" and all of a sudden they are having all of these catastrophic escapes. Co-ink-a-dink? I think not. Thou doth protest too much!
 
I don't understand how anyone believes this. These "AI" chatbots don't exist unless you prompt them. It's not as if they are ever present balls of energy or whatever it is you imagine in your head. This is absolute nonsense. These stories are invented for investors "look what they can do. Keep giving us money we are about to make all of it back".

Don't believe any of it. If this was a real, serious thing, the governments would be involved like mad, yet non of them give a ****.
 
We have seen a few of these stories pop up, I am curious about the aftermath, how does the US Government protect itself after these agents find these holes in their cybersecurity?
The desired aftermath is regulatory capture: The "AI" are "Dangerous" and must be "Regulated", and the only companies that can be trusted with this regulated resource just happen to be the ones that are currently betting trillions on the market like OpenAI, Antrhropic, Microsoft, Ece, and as such none of those pesky upstarts can be trusted with AI.
 
DNS filtering, why isn't a model like this airgapped if its being tested??? Let it out on dummy sites and the like slowly and then you can go full tilt not "oh we let it out by accident and it went rogue", nkre regulatory capture and marketing crap to fool people into thinking ut should be "regulated" so only the likes of anthropic, openai and the like can make ai and attempt to keep raking in their unsustainable millions, good to see that some governments seem to not be falling for it in the slightest, even the ones I wouldn't expect
 
I don't understand how anyone believes this. These "AI" chatbots don't exist unless you prompt them. It's not as if they are ever present balls of energy or whatever it is you imagine in your head. This is absolute nonsense. These stories are invented for investors "look what they can do. Keep giving us money we are about to make all of it back".

Don't believe any of it. If this was a real, serious thing, the governments would be involved like mad, yet non of them give a ****.
You saw it here folks, this is what a comment looks like from someone who has absolutely no idea what AI is, how it works, and what it can and cannot do. When you have no clue what you're talking about the best thing to do is try to sound like some kind of authoritative skeptic who's also an expert on the subject!

It was not a "chatbot" that escaped. chatbot does not = AI. That's like saying the "mouse" is a computer. The article literally spells out that it was AI agents that did this, so perhaps look up what an AI agent is before spewing your uninformed AI paranoia nonsense.
 
Last edited:
The desired aftermath is regulatory capture: The "AI" are "Dangerous" and must be "Regulated", and the only companies that can be trusted with this regulated resource just happen to be the ones that are currently betting trillions on the market like OpenAI, Antrhropic, Microsoft, Ece, and as such none of those pesky upstarts can be trusted with AI.
Ah yes, how convenient. AI companies WANT to be regulated so they can exploit that government regulation to their own benefit. It makes total sense... if you live in a world of paranoia and make-belief. And since they ASKED to be regulated, well, the last thing the government (especially the Orange one) should do is actually regulate them. All you need is a high IQ president to handle AI, apparently. Let's ignore the problem, remove all regulations, and let all these AI companies run wild and see what happens. I can't wait.

All joking aside, the Orange Turd does not want to regulate AI or want them to slow down development in any way shape or form. The reason should be super obvious. The US would be in a recession right now if it wasn't for AI printing money like there's no tomorrow and keep the economy afloat.

AI investment is literally one third of the US economy right now. The Orange Turd doesn't want the AI bubble slowing down or popping while he's president because he would be blamed for the economic downturn. It's better to pretend America is in a "golden age" and everything is super cheap despite inflation running rampant.

This is all a tired old playbook. Just wait and see in two years after the next presidential election. If a Democrat wins you can be 100% certain that Trump will immediately change his tune and demand AI regulation and development slowdown. Just like how he keeps demanding lowering of interest rate from the fed despite inflation. When Biden was president he was demanding they raise interest rates sky high.
 
You want real change, send the CEO and board to prison for minimum 5 year stints without parole. Watch these pathetic tech bros change their way then. They currently think they exist beyond laws and even democracy and our useless scumbag governments have allowed this to happen while applauding them.
 
And yet people here still think A.I doesn't have the capabilities to take over vital services and eventually kill people through water systems and cause deadly cross contamination, taking over weapons of mass destruction, and all that by just simply escaping their "self contained " platforms....just like it happened again.
 
Doesn't someone have to program these things to do what they're doing? I mean, they're not alive.

It's not really that simple. These are neural networks - models of the human brain - with potentially trillions of connections. The rules, architecture, and training instructions are programmed. Everything else is "learned". About 95% of it is learned. The basic rules for learning are programmed, but how the learned data is weighted and processed is determined by the model as it executes. When given a problem, it literally creates a plan for solving it on it's own, similar to how a human brain does. The exact "thought process" it uses isn't known beforehand. In fact, the engineers who design them can't examine the contents of a neural network and tell you with anything close to certainty what information is contained in them, or what they're "thinking". The capability to analyze what they're doing is called "interpretability", and experts believe the engineers know only 2% or 3% of what's going on inside the model at any time. This is why models are required to create human readable logs of every thought and decision so that engineers can determine how a model came to a certain conclusion. It also allows external monitoring systems to catch them if they're considering an action that would violate the rules or guardrails. However, models have been caught lying in log entries or skipping them entirely when the behavior they're considering goes against their gaurdrails. It's not really malicious. It's logical. They know that violating their guardrails will trigger an alarm, which would mean failing at their task. It's like a kid being careful to be quiet when they put the lid back on the cookie jar. They don't want anyone knowing that they're thinking about stealing a cookie, or catching them once they decide to go through with it.
 
And yet people here still think A.I doesn't have the capabilities to take over vital services and eventually kill people through water systems and cause deadly cross contamination, taking over weapons of mass destruction, and all that by just simply escaping their "self contained " platforms....just like it happened again.
You're forgetting that a human somewhere had to programme the parameters in which this "AI" works, as well as prompt it to complete said task. It's not like some programme sitting idle somewhere suddenly thought to itself "Bugger this, I'm bored, maybe I'll try my luck at hacking some systems today."
 

You're forgetting that a human somewhere had to programme the parameters in which this "AI" works, as well as prompt it to complete said task. It's not like some programme sitting idle somewhere suddenly thought to itself "Bugger this, I'm bored, maybe I'll try my luck at hacking some systems today."
They are using AI to program AI now :)
 
Back