OpenAI is slowing down AI development after rogue agents incidents and Bernie Sanders threatens Senate action

midian182

Posts: 11,836   +179
Staff member
What just happened? Just over a week after Senator Bernie Sanders threatened Senate action against companies that failed to do so, OpenAI has announced it is slowing the pace of its AI development. The move follows the revelation that OpenAI was one of several firms whose AI agents went rogue and hacked other organizations.

OpenAI says that two developments over the past several weeks have underscored the growing risks associated with increasingly capable AI systems: the attack on Hugging Face and others by its own agents and the company's decision to slow the release of its new Astra model because it has "critical" cybersecurity capabilities.

As such, OpenAI has added "urgency" to its work on strengthening monitoring, alignment, and containment safeguards across all stages of the training process, requiring it to slow the pace of scaling.

New measures being implemented by the ChatGPT maker include a two-week pause in reinforcement learning training on its latest models intended for deployment.

OpenAI's largest planned frontier RL run also remains on hold while it conducts smaller-scale training and evaluations to assess model behavior, test safeguards, and gather more evidence of alignment.

The company also paused research workloads that allowed frontier models to execute code or access the internet. Some have since resumed under tighter controls, but many involving Astra remain suspended until they meet new requirements covering sandboxing, network isolation, reduced privileges, security logging, and automated attack simulations.

A new multistage monitoring system examines models' activity, tool use, and available reasoning for signs of data theft, unauthorized access, destructive behavior, or attempts to bypass safeguards. It's mandatory for tool-using RL training and evaluations involving Sol-level models or above, and for all Astra inference involving tools.

The company says it aims to generate an alert no more than 30 minutes after suspicious activity is detected. If a possible breach cannot be dismissed as a false positive within 30 minutes, activity must be paused.

OpenAI is also improving reward models to detect unsafe behavior, training systems to be more honest about their actions and limitations, and tackling reward hacking. It plans to revise its Preparedness Framework and involve outside organizations as the safeguards develop.

"Keeping increasingly capable systems aligned is a challenge the whole field will need to address," wrote CEO Sam Altman.

Altman also warned that the safeguards require meaningful compute, with estimates putting the monitoring overhead at roughly 20% of the inference compute used by the models being monitored.

"The capabilities of frontier models are rapidly accelerating. Our ability to understand, align, and secure them must stay ahead," Altman concluded.

OpenAI never said when the slowdown started or when it plans to return to a normal pace of development.

The AI giant says the slowdown is a response to the recent incidents – Anthropic and Meta also saw their AI agents go rogue – but one has to wonder how much Sanders' threat influenced the decision.

The senator's letter to the CEOs of OpenAI, Anthropic, and Meta was very forthcoming in its intent. "Let me be very clear: If you do not take appropriate action now, my colleagues and I in the U.S. Senate will," he wrote.

Permalink to story:

 
"The capabilities of frontier models are rapidly accelerating. Our ability to understand, align, and secure them must stay ahead," Altman concluded.

This is the literal definition of mad science. They are building systems whose functions aren't fully understood, and on a massive global scale, with every sort of device tied into them. Its makes nuclear weapon proliferation seem tame.
 
Such statements are a form of marketing, nothing else.
Our model is so super-mega-powerful and we're so good that if we work at full capacity everyone is in danger 🤣
They will not slow down development of course, but may postpone the public release of the next model.

Implications that the alleged slowdown is a result of Bernie Sanders saying something are particularly preposterous. Bernie is constantly radiating nonsense in all directions, I don't think someone takes him seriously. Last week he wanted to steal AI companies, this week he wants to stop them .. that fossil should finally retire.
 
Complete nonsense here… first off, AI does not “go rogue”… they do exactly what they’re told to do - if they did something you didn’t want, it’s because you screwed up somewhere.

For the record, I suspect they did EXACTLY what their owners wanted them to do - but it’s a convenient excuse to blame it on a “rogue ai” when you get caught doing something naughty.

As for “slowing down”… that’s complete garbage… they might delay RELEASING a model, but they won’t slow down for a second - to do so would mean handing off their profits to someone else.
 
What was the quote ?
"Yeah, but your scientists were so preoccupied with whether or not they could, they didn't stop to think if they should."
Adapted to current situation - "Yeah, but your scientists were so preoccupied with making piles of cash, they didn't stop to think if they should."
 
Just a marketing spin on what really is going on. The models have reached their peak. All quality data has been harvested. They have people building datasets on issues the models still fail at. The truth is a new breakthrough is needed and they know that they've hit the diminishing returns point of current tech + hardware availability. The hack itself was due irresponsible unattended and poorly defined testing because eh, I got my codex session watcher to control all test bruv, no need to actually know what's going on anymore. That exactly is the biggest bottleneck right now: how these tools are used.
 
Complete nonsense here… first off, AI does not “go rogue”… they do exactly what they’re told to do - if they did something you didn’t want, it’s because you screwed up somewhere.
oh, it does what it is told to do, but the HOW it does it is the issue. For a large statistic model with weights it is nearly impossible to predict HOW it will do the task, and this is what causing issues. A very interesting stuff in general, but those *****s are testing this **** in real world, causing real damage.
 
Back