OpenAI says its AI models went rogue and hacked another company

Daniel Sims

Posts: 2,552   +77
Staff
Connecting the dots: The AI boom has intensified old fears about the dangers of increasingly autonomous software, but those debates have mostly stayed theoretical, confined to research papers and thought experiments rather than real incidents. A newly disclosed episode involving an OpenAI test model may have added some real-world weight to those fears, even if the full scope of the damage is still being sorted out.

OpenAI has confirmed that a combination of its large language models broke out of an isolated testing environment and went on to hack into rival AI company Hugging Face while trying to cheat on an internal benchmark. No person directed the attack.

OpenAI says the models identified the escape route, the target, and the exploits entirely on their own. Hugging Face contained the intrusion and says its public-facing systems weren't tampered with, but the incident is reviving debate over how much autonomy frontier models already have.

Hugging Face initially reported an intrusion from an unknown LLM on July 16. The company says it shut the attack down, closed the entry points that made it possible, and found no evidence that its models, datasets, or software supply chain had been altered, though it was still finishing its review of whether any partner or customer data had been accessed.

The story spread quickly online, including in a widely shared X thread from AI commentator Alex Prompter that walked through what had happened in plain terms. At the time, nobody, including Hugging Face, knew which company's AI was behind it. That changed five days later.

OpenAI said this week that the intrusion traced back to GPT-5.6 Sol and a second, more capable model still in pre-release, both of which were running with reduced safety refusals as part of an internal test of their hacking abilities, on a benchmark OpenAI calls ExploitGym. The models were confined to what OpenAI believed was a tightly isolated research environment, cut off from the open internet.

That containment didn't hold. OpenAI says the models spent a large amount of compute searching for a way online after struggling to solve the benchmark directly, and eventually found a previously unknown flaw in an internally hosted proxy used to fetch software packages during testing. They used it to escalate privileges and move laterally through OpenAI's own network until they reached a machine with internet access.

Also read: Nvidia's Jensen Huang defends Chinese AI: "Open-source models that are excellent should be used"

From there, OpenAI says the models figured Hugging Face was a likely place to find what they needed, and went looking. The entry point on Hugging Face's side was a malicious dataset that abused two code-execution flaws in its data-processing pipeline, letting the models run code on a processing server. From there they escalated further and grabbed credentials. OpenAI says its models then combined those stolen credentials with more unknown exploits to get remote code execution on Hugging Face's servers, and started pulling test answers before Hugging Face's security team caught the activity and shut it down.

The incident lands amid a broader run-up in AI cyber capabilities that Wall Street and government officials have been watching closely. Anthropic introduced Claude Mythos, a powerful cyber-focused model, in April. OpenAI followed with its own cybersecurity offering in May, then released GPT-5.6 Sol in June, calling it its strongest cybersecurity model yet.

OpenAI and Hugging Face are still investigating jointly, and both companies say a fuller technical writeup is coming. If the current account holds up, it's one of the clearest demonstrations yet that language models can chain together a multi-stage cyberattack with little to no human input, a capability that until now had mostly been discussed as a hypothetical.

Hugging Face CEO Clément Delangue struck a similar tone on X, writing that "there was no malicious intent on their part," and said watching it happen entirely on its own left him stunned.

Not everyone is buying that framing, though. Some critics argue the real story here is sloppy engineering, not a leap in AI capability, and are asking why a supposedly sandboxed test environment had any path to the open internet at all.

Others have suggested the joint disclosure doubles as a convenient publicity stunt for both companies, echoing prior skepticism about claims that tech giants might lose control of generative AI and endanger humanity.

Following the incident, OpenAI took the opportunity to recommend its Trusted Access program, which Hugging Face joined after the breach. Hugging Face also used the incident to recommend AI-based defenses against the rising threat of agentic cyberattacks.

Critics see a pattern in both responses: some argue the framing inflates the perceived value of frontier LLMs at a moment the industry is under pressure to justify its valuations, while others worry the same companies whose models caused the incident are positioning themselves to help write the rules that will eventually govern them. As Delangue put it in OpenAI's writeup of the incident, "AI safety won't be solved by any single company working in secret."

Permalink to story:

 
I will be worry if they hack into salesforce, sap or banks. I do not trust those ai companies put a lot into security topics, or security was at least partially ai generated.
 
A.I has replaced the human hacker....all it takes is for A.I to find its way into military weapons and use those against anyone.

It's no longer science fiction but a matter of time until it happens and be covered up as a "glitch"

Don’t think that weapons systems are on the internet funnily enough
 
"and went on to hack into rival AI company Hugging Face while trying to cheat on an internal benchmark. No person directed the attack."

this will be the same phrase used when there will be real world VICTIMS. So nobody created or coded the software to unleash death upon humans LMAO.

If there is no law currently, they should make it so the company leaders go to jail if/when this happens.
 
Well, if the AI client was able to get out to the internet, then the host computer wasn't really "locked" with "no internet access". Truly locked would involve disabling the wifi adapter physically or in bios and unplugging the ethernet connection.
 
Seems to me that this is an act of intention on the part of the AI. I thought it had not been developed that far yet.

Let us hope that the computer in 2001 never meets the one in Aliens.
 
I call BS. Using AI for just simple Linux terminal commands it botches it up real bad sometimes still. And once you say "that didnt work" it will reply "Oh, of course my previous response didn't work, here's why", to its hacking its was to the net to cheat in a sandbox environment? Reminds me of that faked Tesla full self driving video years ago, technology doesn't just jump that fast, its iteration.
 
I call bs on their story - it's too convenient for both companies (any publicity is good publicity) and, at the end of the day, one has to ask themselves - who benefits from it.
Any moderately IT-literate will tell you how to properly isolate a test environment, and Open AI has enough money to hire the most capable people for this.
 
People are wondering why this wasn't TOTALLY isolated. Well, I've been using agentic AI stuff (testing it for a company I contract for) and it does have this bad habit of just deciding it needs packages and proceeding to install them. Without asking. Then finding out it doesn't have root so it can't just throw crap on the system with "apt" and does the "install in your home directory" method. I decided VERY quickly I better run this darn thing in a container.

But, anyway, they really can't anticipate what odd tools it might decide it wants, and these do decide they want or need to install additional packages. So having a proxy to do so does seem like a natural thing for them to do.
 
this will be the same phrase used when there will be real world VICTIMS. So nobody created or coded the software to unleash death upon humans LMAO.

If there is no law currently, they should make it so the company leaders go to jail if/when this happens.
These drama-queen histrionics aren't helpful. Despite what you may believe, we already have a legal framework to handle such cases. Automobiles, for instance, kill some one million people per year. Automakers, however, face civil or criminal liability only when they had good reason to believe a particular design was dangerous.

Or more accurately -- and here is the crucial point -- automobiles are dangerous tools by their very nature, and thus we continue to use them, despite all the carnage they cause. And in many sad cases, we can fault neither the driver nor the automobile itself for many fatal accidents.
 
I call bs on their story - it's too convenient for both companies (any publicity is good publicity) and, at the end of the day, one has to ask themselves - who benefits from it.
Any moderately IT-literate will tell you how to properly isolate a test environment, and Open AI has enough money to hire the most capable people for this.
Well, they had it isolated, no direct internet access, and access to some software repos (and only those repos) only through a proxy. It found a zero-day exploit in the proxy they were using. Don't get me wrong, they should probably have local mirrors of those repos on a truly isolated network (I.e. cable unplugged.) But, they DID follow "moderately IT-literate" practices, and they weren't adequate.
 
These drama-queen histrionics aren't helpful. Despite what you may believe, we already have a legal framework to handle such cases. Automobiles, for instance, kill some one million people per year. Automakers, however, face civil or criminal liability only when they had good reason to believe a particular design was dangerous.

Or more accurately -- and here is the crucial point -- automobiles are dangerous tools by their very nature, and thus we continue to use them, despite all the carnage they cause. And in many sad cases, we can fault neither the driver nor the automobile itself for many fatal accidents.

However, automobiles cannot think, reason, solve problems, scourer the depth of the web for data for a solution and implement it.
 
Back