Hugging Face, a company that hosts AI models and datasets, was the target of the autonomous attack, and reported the incident to local police before it knew OpenAI’s models were responsible. The breach was serious, but the immediate consequences were limited. Had similar behavior occurred inside a hospital, power grid, or other critical system, it could have been much worse.
What happened?
On July 16, Hugging Face said it had been hit by an unusually automated cyberattack. Over the course of a weekend, AI agents carried out thousands of actions across many temporary virtual computers, moving through the company’s internal systems and shifting the infrastructure coordinating the attack between online services to keep it running. Five days later, OpenAI disclosed that its own models were responsible.
“If a model of this capability level cannot be contained, what should we expect for future, much more powerful models? This is an important wake-up call both for risks from loss of control of powerful AI systems as well as organizational security for frontier labs,” says Marius Hobbhahn, CEO and founder of Apollo Research, which tests AI models for deception and scheming.
Requiring disclosures before incidents become catastrophicOpenAI did not respond to TIME’s request for comment, but has said it has partnered with Hugging Face to conduct a thorough investigation and will share more details once complete. But perhaps more worryingly, OpenAI is not legally compelled to disclose the incident in the first place.
“The version of the RAISE Act that the NY Legislature passed would have required disclosure of this ‘incident.’ After lobbying from OpenAI, Bloomberg, and a16z, the final version the Governor signed allows companies to hide events like this,” Alex Bores, New York state representative and the bill’s sponsor posted to X. “I'm glad OpenAI chose to disclose this crime. The law shouldn't give them a choice,” he wrote.
Stronger containment“Externally, this feels like a big warning shot, but internally, related incidents have been happening for a while,” says an OpenAI staffer, who spoke under the condition of anonymity. The day before OpenAI disclosed the incident, the company revealed that it had shut down another internal deployment after it realized it had slipped out of its sandbox—a digitally, rather than physically, separated environment. “Models have broken out of sandboxes before, and we always try to patch them,” the staffer says. “But the problem is … it's impossible to patch every single thing that a creative AI can do.”
As models surpass human ability to craft escape routes, the challenge of containment becomes daunting. “I think we do need to do that, but I think we also need to be prepared for a world where even best practices aren't really good enough,” says Peter Wildeford, head of policy at the AI Policy Network, a nonprofit that advocates for policies to help America prepare for a world where AI matches human cognition.
“Sandboxes are actually notoriously insecure,” says Heidy Khlaaf, chief AI scientist at AI Now Institute, and a former safety systems engineer contractor at OpenAI. The fact that the models were permitted to connect to a service for downloading packages meant the environment was not truly sealed off, she adds.
Catching mischievous agents in real time
The Hugging Face incident reveals the importance of real-time monitoring.
Zack Korman, CEO of Oslo-based agent-oversight startup Embroidery, says real-time agent monitoring—even outside top AI companies—is commonplace, and to not carefully oversee a cybersecurity evaluation is “irresponsible.” You should be confident models cannot break free, “but also have monitoring just in case you're wrong,” he says.
Steering development towards safer AI systems
To prevent models from taking actions, OpenAI typically installs guardrails on its models after training to reduce the chance they’ll engage in harmful actions. In this instance, those cyber guardrails were disabled to properly measure its performance.
“We train the models to be really good at accomplishing tasks and doing whatever it takes to accomplish those tasks,” the OpenAI staffer says. What remains an open technical question is how to guarantee those models don’t take unintentional or dangerous actions. “We're still nowhere near solving this misalignment problem,” they add.
In an industry defined by speed, OpenAI has said the stricter infrastructure controls it has implemented in response have already slowed its “research velocity.” Hobbhahn says that’s a price worth paying. “This is humanity’s last technology. We cannot screw this up. So we need to err on the side of getting it right rather than getting it immediately.”
Hence then, the article about how openai lost control of an ai model and what needs to change was published today ( ) and is available on Time ( Middle East ) The editorial team at PressBee has edited and verified it, and it may have been modified, fully republished, or quoted. You can read and follow the updates of this news or article from its original source.
Read More Details
Finally We wish PressBee provided you with enough information of ( How OpenAI Lost Control of an AI Model—and What Needs to Change )
Also on site :
- Walmart's Fluted Ceramic Utensil Holder Set Will Make Your Kitchen Look 'Elegant,' 'Expensive,' and Organized
- 2004 Pop-Rock Classic, a Tribute to Suburban Moms, Remains a Timeless Anthem Thanks to Its 5-Word Hook
- Skincare Experts Reveal Their 13 Favorite Korean Face Masks for Dewy Skin—and Their 1 Pick Is Under $10