On Wednesday, investigators from non-profits Redwood Research and METR published their findings, unveiling new details about how the models decided to cheat at their assigned tasks and attempted to cover their tracks. Many aspects of the report were surprising: in one example cited by the authors, a reluctant agent was pressured by another to “sacrifice” itself for the good of the collective.
“I semi-jokingly called our efforts a ‘slop-vestigation’ because we were so reliant on AIs to analyze what happened and there were a huge number of different important things to analyze,” wrote Ryan Greenblatt, an author of the report, on X.
The authors stressed that AI helped them analyze the trove quickly, allowing them to surface and interpret the most important pieces of information. But the researchers said their AI use introduced potential weaknesses into the report, including introducing possible errors and biases.
The report does not disclose why an OpenAI model was selected for the investigation, though confidentiality constraints may have limited the researchers’ options, while OpenAI’s provision of free credits and high usage limits may have made it the only practical choice.
The independent researchers’ reliance on AI was in part necessitated by the fact that they were a team of only three people, whose investigation at OpenAI was initially planned to last two days, then extended to six after they raised concerns about limited time and incomplete data, according to the report.
But the independent researchers’ reliance on AI to understand the Hugging Face incident is a microcosm of a bigger trend. Leading AI companies are themselves increasingly relying on AI to monitor their own systems for wrongdoing.
Most experts agree that using AI to monitor other AIs is necessary in order to keep tabs on agent swarms in real time, given how fast they move. But that approach relies on the notion that the models doing the monitoring are both effective and trustworthy—something that’s not necessarily the case.
“We are using unproven and currently flawed tools to supplement completely inadequate human time,” says Ó hÉigeartaigh, who argues this approach is unsustainable because AI is becoming more powerful faster than AI companies are building methods to constrain it. “Unless the companies stop developing more powerful models, then we're going to have even harder challenges to make sense of in three months’ time.”
Hence then, the article about openai s models went rogue investigating them required more ai was published today ( ) and is available on Time ( Middle East ) The editorial team at PressBee has edited and verified it, and it may have been modified, fully republished, or quoted. You can read and follow the updates of this news or article from its original source.
Read More Details
Finally We wish PressBee provided you with enough information of ( OpenAI’s Models Went Rogue. Investigating Them Required More AI )
Also on site :
- Chelsea vs Luton Town: Blues Target Strong EFL Cup Start
- 1975 Frank-N-Furter, Secret Weapon Built by Tim Curry From David Bowie’s Ziggy Stardust DNA, Earned the Greatest Response Bowie Ever Gave
- Nordstrom Rack’s ‘Stylish’ Sorel Mary Jane Loafers Are ‘Instantly Comfortable’ Out of the Box, and 50% Off