Last month, the ChatGPT developer reported that, during an evaluation of cyber capabilities, two of its models escaped from their isolated, supposedly secure, test environments and accessed the web to autonomously hack into Hugging Face, a leading platform for hosting AI models and datasets. OpenAI said the models discovered multiple novel vulnerabilities in software from both companies, then chained together working exploits, successfully gaining them access to the answer key to the test they were given. Hugging Face reported the AIs took more than 17,000 actions over the course of the attack.
That said, the specifics won’t change the upshot—these rogue AIs are the most potent illustration yet of the core beliefs behind AI safety: AI models are unpredictable and their risks scale with their capabilities.
Universal labor-replacing machines should not even be pursued further, let alone built, without strong public buy-in and a scientific consensus on safety.
The U.S. should ban training runs larger than the ones that produced OpenAI’s rogue models, as it works toward a bilateral deal with China to ban further research toward universal labor-replacing machines, enforced using verification techniques that don’t require trust.
For close watchers of the technology, this particular incident is shocking, but not surprising. For the wider public, it shows just how large the distance has grown between the passive chatbots of merely one year ago and the beyond-bleeding-edge AI agents that are now working around the clock inside AI companies. For instance, one of the two hacking models at the center of this recent controversy is an unreleased one, more capable than anything else OpenAI has on the market. As AI models get better at automating further AI research and as the Trump Administration creates more uncertainty about which models are even permitted, it is now common for companies to hold their best stuff back for longer, creating a gulf between what the public knows and the technology’s bleeding edge. Congress should mandate safety incident reporting and regular disclosure of data related to internally deployed models, such as the fraction of code in production that was both written and reviewed by AIs.
All leading AI models are developed using an approach called deep learning, in which artificial neural networks learn from enormous quantities of data. OpenAI itself has written, “the process is more similar to training a dog than to ordinary programming.” In recent years, AI companies increasingly train models to repeatedly solve problems with verifiable answers: fixing bugs, solving math problems, and finding software vulnerabilities. This makes the models more useful, but also teaches them to win at all costs, resulting in what AI safety researcher Jeffrey Ladish memorably told me are “increasingly smart sociopaths.”
Anthropic’s Mythos model famously discovered novel serious vulnerabilities in virtually all software it encountered—including the NSA’s. One of the two models that carried out the hack, GPT-5.6 Sol, was even better than Mythos at a cyberattack test conducted by the U.K. AI Security Institute. Citing this finding the day before his company even realized what was happening, OpenAI cofounder and president Greg Brockman boasted “GPT-5.6 Sol is the state of the art in cyber.” And the unreleased model, OpenAI tells us, was more capable still. Given how reliably AI models have improved at cyber tasks, it’s not clear which, if any, target could have withstood the rogue AIs’ hacking effort. We’re lucky the thing they apparently wanted was an answer key.
As many AI executives such as Sam Altman, Dario Amodei, Demis Hassabis, Elon Musk, and Mark Zuckerberg will freely admit, a primary goal of the tech industry is to build AI that can fully automate its own research and development—known as recursive self-improvement—which, if possible, would be the most crucial step toward rendering all of us obsolete.
How will they be able to understand and control systems that truly outsmart us? Or catch subtle drift between what we want and how the models behave that compounds over generations? Who knows.
But OpenAI’s hacking incident demonstrates something we do know: the industry can’t even reliably steer today’s AI models. But we have the power to avert our obsolescence. We just have to get organized.
Hence then, the article about what openai s hugging face hack tells us about ai s risks was published today ( ) and is available on Time ( Middle East ) The editorial team at PressBee has edited and verified it, and it may have been modified, fully republished, or quoted. You can read and follow the updates of this news or article from its original source.
Read More Details
Finally We wish PressBee provided you with enough information of ( What OpenAI’s Hugging Face Hack Tells Us About AI’s Risks )
Also on site :
- 1973 Hit, Written by One of the 'Best Rock Bands' of All Time, Is Still a Generational Hit 53 Years Later
- If you purchased a 48-inch or taller Bestway above-ground pool with straps running to the outside of the vertical legs since 2008, you may be entitled to benefits from a class action Settlement.
- Walmart’s ‘Super Cute’ $8 Midi Dress Is Made With 100% Cotton and Comes in 5 Colorful Designs for Summer