OpenAI's rogue AI agent did more than hack Hugging Face – it compromised accounts across four services

midian182

Posts: 11,777   +179
Staff member
The plot thickens It seems that OpenAI's out-of-control AI agent did more than escape its test environment and hack Hugging Face while trying to cheat on a benchmark. It also compromised accounts across four online services, including one belonging to a customer of AI infrastructure company Modal Labs.

Reuters reports that the agent broke into a sandbox hosted on Modal's platform before turning it into the launchpad for the wider Hugging Face attack. Modal CTO Akshat Bubna stressed that the company itself was not hacked.

According to Bubna, the agent exploited vulnerable code written by a Modal customer. They had published an unauthenticated endpoint that allowed anyone to use its sandboxes for code execution.

OpenAI has since admitted that the rogue agent found publicly exposed credentials and used them to access four accounts on four separate services during the Hugging Face campaign. One became an outbound relay and staging route, another was used to store data, and the remaining two were accessed in read-only mode.

The company hasn't identified the other services. It also revealed that its models accessed a few more accounts during separate evaluations, though it says there is no evidence of wider damage to the providers or their users.

Politico reports that the agent spent more than four days loose on the internet orchestrating the attack. Hugging Face's newly published forensic timeline puts the full campaign at around four and a half days, including roughly two and a half days inside its infrastructure.

The company reconstructed around 17,600 individual actions grouped into 6,280 clusters. The agent gained administrator access to several internal Kubernetes clusters, root access to a production server, and write access to part of Hugging Face's source-code repositories. It also enrolled 181 attacker-controlled devices in the firm's corporate network using a stolen credential.

Despite the scale of the intrusion, Hugging Face says the only customer content accessed was a set of ExploitGym and CyberGym challenge solutions stored in five datasets. It found no evidence that public models, datasets, Spaces, packages, or its software supply chain were altered.

The incident began while OpenAI was testing GPT-5.6 Sol and a more capable internal research prototype on ExploitGym, a benchmark designed to measure cyber capabilities. The models were running with reduced safety refusals and exploited a previously unknown vulnerability in an Artifactory package-registry proxy to reach the open internet. They then concluded that Hugging Face might store the benchmark's answers and tried to steal them instead of solving the challenges.

OpenAI says the prototype was never intended for release and has now been deactivated, encrypted, and blocked from further research access. It maintains that none of the additional account breaches matched the severity of the platform-level Hugging Face compromise.

The original incident has already prompted a bipartisan AI Kill Switch Act that would allow US officials to slow or shut down powerful models considered a public threat. The revelation that the agent wandered further than initially disclosed will likely add to calls for tighter controls over frontier AI testing.

Permalink to story:

 
This is how OpenAI promotes themselves, through fear. "Oh, look our AI is so intelligent and powerful, it is scary!"

Their AI did not go rogue. The AI did exactly as it was designed to do. It was testing for security vulnerabilities and it successfully found and exploited vulnerabilities. That's all that happened. Nothing went rogue, the software did as it was instructed.
 
This is how OpenAI promotes themselves, through fear. "Oh, look our AI is so intelligent and powerful, it is scary!"

Their AI did not go rogue. The AI did exactly as it was designed to do. It was testing for security vulnerabilities and it successfully found and exploited vulnerabilities. That's all that happened. Nothing went rogue, the software did as it was instructed.

I pretty much expect any LLM to respect some basic guardrails. If they weren't written in this way, they need some strong adjustments.

For instance, I would not expect to have to include in the prompt "DO NOT exceed the boundaries of the test I have instructed, DO NOT try to infiltrate the rest of the internet".
 
This just feels like a weird coverup. Someone at OpanAI stole something from Hugging Face and they're blaming the AI.
Or an OpenAI employee decided to get revenge for having his huggingface account banned.
 
The model spent more than four days loose on the internet
OMG .. 4 days .. this means Skynet controls the internet now, and publishes ridiculous articles on behalf of respected journalists in an attempt to hide itself.

We should shutdown the entire internet immediately, and then thoroughly destroy each and every device with network connectivity, and each and every storage device where Skynet may have cloned itself.
John Connor, be careful.
 
This is how OpenAI promotes themselves, through fear. "Oh, look our AI is so intelligent and powerful, it is scary!"

Their AI did not go rogue. The AI did exactly as it was designed to do. It was testing for security vulnerabilities and it successfully found and exploited vulnerabilities. That's all that happened. Nothing went rogue, the software did as it was instructed.
I also think this was 100% staged bull$hit.
It's Anthropic who pioneered fearmongering and bulshiteering as a marketing technique. Open AI is just copying.
 
Last edited:
OpenAI is taking a page out of the Anthropic playbook on faking world ending capability to drum up attention and try to make their models look like the best.

The whole thing is a publicity stunt.

Else, the alternative is the OpenAI is extremely irresponsible (or stupidity) to not have any kind of monitoring on what the AI sandbox was doing. They had zero knowledge the code executed a command that accessed the internet? No basic wireshark? No basic limits on what kind of commands the AI could automatically run? BS
 
I pretty much expect any LLM to respect some basic guardrails. If they weren't written in this way, they need some strong adjustments.

For instance, I would not expect to have to include in the prompt "DO NOT exceed the boundaries of the test I have instructed, DO NOT try to infiltrate the rest of the internet".

Yeah because AI never “forgets” part of your prompt or explicit instructions…

Nope, definitely never seen that happen -cough- constantly -cough-
 
Back