The unauthorised incursion into the inner workings of Hugging Face, an AI model library and hosting platform, by an unknown model was concerning enough when it emerged last week. It became more alarming when OpenAI, the makers of ChatGPT, admitted the intruder had been one of its own models which had gone rogue.
OpenAI had asked some of its most capable models to complete a cyber security test called ExploitGym. Rather than solve the challenge as intended, the agent searched for another route to a high score.
According to OpenAI’s account, the model discovered a previously unknown flaw in software controlling its test environment or “sandbox”, and then escaped on to the open internet, stole credentials and broke into Hugging Face, which hosts AI models and datasets. Once it entered Hugging Face’s systems, it sought out the answers to the test.
The models, which include GPT-5.6 Sol, available to current ChatGPT subscribers, had their usual safety provisions reduced for the evaluation. To try and limit risks, OpenAI kept the agent in a specific test environment – but the agent appears to have breached its boundaries, and wreaked havoc in what OpenAI called an “unprecedented cyber incident”. The victim, Hugging Face, tracked the agent making more than 17,000 different actions on its systems.
‘Nobody told it not to cheat’
But what should we make of the AI activity, and does it show the technology has a twisted sense of morality?
“It’s not immoral. It’s just amoral,” said Alan Woodward, visiting professor of computer science at the University of Surrey. “Nobody told it not to cheat. You just told it to pass the test.”
While a human employee might think twice about staying within the confines of an experiment, the AI model thought nothing of breaking out, stealing passwords and intruding into another company’s systems.
“They don’t have emotional guilt and they don’t have legal liability,” says Noah Giansiracusa, associate professor of mathematics at Bentley University. “So why wouldn’t they find all kinds of ways of lying and cheating and stealing and doing whatever it takes?”
While OpenAI called it unprecedented, it’s not unusual. The UK’s AI Security Institute revealed that an AI it was testing went rogue and tried hacking its system. The institute also revealed every leading model from big companies tried to cheat the set test in order to complete the task using unauthorised means.
What is unprecedented is that each of those episodes happened in controlled internal experiments. The Hugging Face breach broke out into a real third party’s live systems, showing that just a security test could spill out and create real-world effects.
And that should concern all of us. AI agents – which can take actions in the digital world rather than simply replying to user queries or requests – are increasingly being connected to companies’ email, software, payments and infrastructure. A system told to restore a bank’s service or keep trains running could cause damage if it discovers that the quickest path to its target involves disabling a safeguard or attacking another system.
Woodward’s worst-case scenario is an agent turning its cyber capabilities against “air traffic control systems or transport systems or water chemical plants”. Such an incident could happen inadvertently because the AI might simply fail to understand the difference between a simulation and real life.
Jake Moore, global cyber security adviser at Eset, doesn’t think this ushers in an era of all-powerful automated cyber attacks. But he says similar incidents are likely as models improve and companies race to deploy them. “AI has just made everything quicker,” he says. “When it’s given a goal, it might keep going, and unfortunately, until it finds what it’s been asked to do.”
So what can we do about it?
For the general public, and those running those systems, stronger cyber security including stronger passwords and two-factor authentication at an individual level. Organisations will want to more tightly limit permissions and design their systems on the assumption that an agent end up attacking them, even inadvertently. More powerful agents should not be able to act unsupervised, Woodward argues.
We can also put AI into the fight against AI. Hugging Face used the Chinese GLM-5.2, similar to Claude or ChatGPT, to analyse and ward off the OpenAI attack because safety restrictions on leading American models blocked them from responding to malicious code. Machine-speed attacks require machine-speed defence.
That makes Donald Trump’s approach to those Chinese models particularly awkward. His White House says it is monitoring the incident, while bipartisan lawmakers have proposed independent audits and an “AI kill switch”. But those close to Trump are reportedly considering limiting access to Chinese AI models in the United States after a newly released one, Kimi K3, reached near state-of-the-art performance. Doing so could stop access to the only defensive capability to kill off the next major attack.
“You’ve got to accept that it’s going to happen,” says Woodward, “and therefore you’ve got to equip those that might be targets with the same type of technology that can respond on the same timescales.”
The future of cyber attacks will soon move from humans versus hackers, to AI versus AI.
Hence then, the article about conniving ai is starting to slip human control we should all be worried was published today ( ) and is available on inews ( Middle East ) The editorial team at PressBee has edited and verified it, and it may have been modified, fully republished, or quoted. You can read and follow the updates of this news or article from its original source.
Read More Details
Finally We wish PressBee provided you with enough information of ( Conniving AI is starting to slip human control. We should all be worried )
Also on site :
- Trump bombs in WHCD return and then unleashes on CNN’s Kaitlan Collins and media in jaw dropping hour-long rant
- El afán de Trump por los aranceles nunca disminuyó. Sus próximos movimientos podrían reconfigurar el comercio
- Trump’s biggest laugh at the WHCA dinner: Making fun of his own health secretary RFK Jr for eating roadkill