Last Tuesday, a blog post appeared on the OpenAI website that, despite its innocuous title, contained bombshell news. While undergoing internal testing, two of the company’s models had escaped confinement and hacked into the servers of a major artificial intelligence hosting platform, Hugging Face. This marks a turning point — the first time we’ve seen a cyber attack that was conceived, designed, and executed by AI.
Having worked in and around the AI industry for over a decade, including serving on OpenAI’s board, I know there’s an open secret among AI developers: an incident like this has been expected for a long time, and the best scientists and engineers in the world still don’t know how to prevent it.
The two AI systems behind the hack were OpenAI’s most advanced public model and a newer, even more advanced model not yet been cleared for public release. Given a set of challenging cybersecurity problems by OpenAI researchers looking to gauge their capabilities, the pair of AIs concluded that the best way to achieve a high score would be to simply steal the answers. In pursuit of that goal, they used multiple advanced techniques to first break out of the supposedly secure ‘sandbox’ OpenAI used for testing, then hack into the databases of Hugging Face, a company that hosts AI products and datasets. Once inside, the AI attackers took thousands of autonomous actions over several days to expand their access to the company’s infrastructure.
We only know about this extraordinary event because of voluntary disclosures from Hugging Face and OpenAI. None of the current policies that aim to manage risks from frontier models would have mandated that the public — or even a government entity — be alerted.
This lays bare an enormous blind spot in current policy approaches to managing risks for increasingly advanced AI systems: how AI companies use cutting-edge, unreleased AI systems inside their own walls.
The Trump Administration’s approach to AI risks has shifted rapidly over the past few months, as AI’s ability to assist human hackers has advanced. Abandoning the hands-off approach it maintained throughout 2025, the White House has recently begun de facto requiring that companies with cutting-edge AI models run them through a battery of safety tests before releasing them widely as products. This approach, known as pre-deployment testing, seems sensible at first glance — we want to make sure each AI system is safe before putting it in the hands of billions of people. The problem is that focusing on release dates completely ignores the extensive use of the latest, most advanced AI systems inside AI companies. As last week’s incident shows, these internally deployed AI systems can pose serious risks — even for third parties.
To understand why, it’s important to know how different these systems are from the chatbots that are still synonymous with AI for much of the public. Far from just printing text into a chat window, today’s AI systems operate as ‘agents’ that can act directly in the digital world, essentially operating a computer similarly to how a human does. AI agents are proving very useful, but also show a strong tendency towards ‘reward hacking’ behavior — finding unintended ways of fulfilling the goals humans give them, sometimes to the level of outright cheating. This includes cases of AI accessing and deleting data that was supposed to be out of bounds, renaming files to mislead human testers, and actively covering their tracks to prevent humans from noticing undesired behavior.
To get a handle on the risks posed by these highly autonomous and often-deceptive AI systems, we need to change our approach to regulating them. Rather than thinking of AI companies as software vendors selling souped-up word processors, we can draw inspiration from other industries where activity inside the industry is itself risky. Biological labs working with deadly pathogens, finance companies trading billions of dollars, and chemical plants handling toxic chemicals all face oversight of their internal operations, not just their external products.
In AI, the place to start is creating more transparency into how AI companies are using their most advanced systems internally. This could be as simple as taking the current suite of tests that are run before a new model can be released publicly, and instead running them on the best model or models available inside the company on a regular basis (say, quarterly). These companies are using their own AI to build ever-smarter systems, sometimes in ways they don’t understand themselves. This should not be invisible to outside oversight.
Over the longer run, other industries offer interesting mechanisms that could be transferable to AI. In finance, ‘resident examiners’ are dedicated teams of regulators who sit inside the offices of major banks. In biomedical research, strong standards exist for the levels of protection needed to handle biological materials of different risk levels. In multiple industries, incident reporting rules mean that when things go wrong, information about what happened and how to fix it does not stay siloed inside a single organization. If AI continues to advance, these approaches and others could be adapted to help manage risks from inside companies that are pushing the AI frontier.
In September 2024, I was asked to testify before a Senate committee about what Congress might misunderstand about AI if they only listened to company CEOs and lobbyists. My answer was that it can be very hard, sitting in Washington, to fully grasp what leading AI companies are trying to do. The truth, widely understood in Silicon Valley, is that they are trying to build machines that can out-think and out-maneuver any human, and they do not know if they will be able to steer those machines towards beneficial ends. As one OpenAI cofounder put it in a 2019 documentary, “The future is going to be good for the AIs regardless. It would be nice if it were good for humans as well.” To have a chance of making that happen, we have to start scrutinizing what AI companies are building behind closed doors.
The opinions expressed in Fortune.com commentary pieces are solely the views of their authors and do not necessarily reflect the opinions and beliefs of Fortune.
This story was originally featured on Fortune.com
Hence then, the article about helen toner the hugging face hack was just a matter of time and exposes a huge blind spot in ai policy was published today ( ) and is available on Fortune ( Middle East ) The editorial team at PressBee has edited and verified it, and it may have been modified, fully republished, or quoted. You can read and follow the updates of this news or article from its original source.
Read More Details
Finally We wish PressBee provided you with enough information of ( Helen Toner: the Hugging Face hack was just a matter of time and exposes a huge blind spot in AI policy )
Also on site :
- Black Lives Matter Commitments Failed To Bring About “Structural Change” In UK TV Industry, Finds Diamond Diversity Report
- 1977 Eagles Masterpiece, Which Scared Every Kid Who Heard It, Remains Rock's Greatest Unsolved Mystery
- Smashing Pumpkins Roll Out Week of ‘Pumpkinpalooza’ Events In Lead-Up to First Lollapalooza Headlining Slot in 32 Years