The models found a previously unknown vulnerability in the infrastructure meant to contain them, gained access to the public internet and broke into Hugging Face, a major platform for hosting AI models and datasets. Their objective, however, was less sinister than the sequence of events might suggest: They were looking for information that would help them complete the cybersecurity test OpenAI had given them.
Independent experts who spoke with Live Science agree that what happened is significant — but they cautioned against interpreting it as an AI system suddenly developing a malicious agenda. The models appear to have pursued the task OpenAI gave them, finding a route to success that their creators had failed to anticipate or adequately block.
OpenAI was testing GPT-5.6 Sol and a more powerful unreleased model using ExploitGym, a benchmark that challenges AI systems to find and exploit software vulnerabilities. The company removed some cybersecurity safeguards that would normally prevent potentially dangerous actions while relying on an isolated environment to keep the models away from the wider internet.
Hugging Face became a target because the models identified it as a possible source of information that could help them complete the ExploitGym challenges. OpenAI said at least one attack chain involved stolen credentials and previously unknown vulnerabilities that eventually enabled the models to execute remote code on Hugging Face systems and access test solutions stored in a production database.
Rather than harboring any malicious intent, the AI models simply wanted to find out more information so they could complete their task. (Image credit: wildpixel/ Getty Images)
Did the AI really "escape"?
It's notable that the models found a flaw in the infrastructure designed to contain an AI and used it to reach the public internet. Describing the models as having "gone rogue," however, risks assigning them unsupported motivations, Buckley said.
Buckley compared it to asking a dog to fetch a ball while leaving the garden gate open. "If the easiest ball for it to find is in the park down the road, that's where it'll head," he said. "You wouldn't say the dog had gone rogue; you'd just say you underestimated how literally it would pursue the task."
What matters more than the models' supposed motives is what they managed to accomplish while pursuing their assigned task.
The lesson isn't that AI has become malicious. Instead, it's that increasingly capable systems will exploit opportunities that humans fail to anticipate.
Katerina Mitrokotsa, a professor of cybersecurity and applied cryptography at the University of St. Gallen in Switzerland, said the containment failure is particularly concerning because another company ultimately paid the price.
OpenAI representatives said they have tightened the infrastructure used for these evaluations. But Mitrokotsa warned that containment becomes harder to guarantee as models improve at performing exactly the kind of exploitation OpenAI was testing.
An AI warning — and an impressive product demonstration
Buckley said announcements from frontier AI companies like OpenAI or Anthropic should be viewed in the context of an industry competing to build ever-more-powerful models.
These companies have every incentive to show both that their models are extraordinarily capable and that they are taking the risks seriously, he added. The Hugging Face incident demonstrates both that OpenAI's models carried out a complex series of operations with considerable autonomy and that its security measures failed to keep them inside the experiment.
Related stories
There are 32 different ways AI can go rogue, scientists say — from hallucinating answers to a complete misalignment with humanityAI self-replication hacks 'no longer purely theoretical,' study finds — but experts say it's too soon to panic'You can't patch your way out of it': Cheap AI worm can spread between devices without human guidance — but how did scientists create it?
The episode, the experts said, leaves OpenAI with a result that is impressive and uncomfortable in equal measure. Its models found previously unknown vulnerabilities and continued pursuing their goal well beyond the boundaries their creators expected, but none of that requires them to have developed malign intentions.
In this incident, OpenAI's new models were given a hacking challenge and they were rewarded for finding a way to solve it. The humans running the experiment simply hadn't anticipated quite how far they might go.
Hence then, the article about no openai s models didn t go rogue when they broke into hugging face here s what really happened was published today ( ) and is available on Live Science ( Middle East ) The editorial team at PressBee has edited and verified it, and it may have been modified, fully republished, or quoted. You can read and follow the updates of this news or article from its original source.
Read More Details
Finally We wish PressBee provided you with enough information of ( No, OpenAI's models didn't go 'rogue' when they broke into Hugging Face. Here's what really happened. )
Also on site :
- NAVER Partners with Brookfield and NVIDIA to Expand Korea's National AI Factory Infrastructure Buildout
- Trump fracasó estrepitosamente en la cena de corresponsales de la Casa Blanca, pero dejó claras sus intenciones
- Audiencia preliminar del cantante d4vd arroja estremecedores detalles sobre la investigación de la muerte de una adolescente