Amid revelations that safety concerns have forced the company to abandon the planned rollout of GPT-6.1 Astra, independent researchers and benchmark creators warn that GPT-6 Astra's record-breaking scores rely on heavy prompt optimization rather than true general intelligence. This comes as experts warn about the risks of future AI systems and AI executives jointly call for a slowdown in AI research.
The benchmarks scores are impressive, but even the creator of one of the most important suggests acing it doesn't automatically mean we've reached AGI. (Image credit: Cheng Xin via Getty Images)
As part of the testing-and-verification process for the new model, OpenAI released benchmark results across a range of tests that measure the capabilities of AI models at various tasks. OpenAI representatives claimed in a statement that the company's new model delivers "state-of-the-art" performance across a variety of fields, including software engineering, the autonomous use of computer programs, mathematics problems, and even scientific research.
"In our early testing, Astra stood out by approaching legal work the way a discerning lawyer does: it distinguishes documents from established records, surfaces unsupported assumptions, and converts gaps into concrete drafting positions," said Niko Grupen, head of applied research at legal services AI developer Harvey as part of OpenAI’s announcement.
"Not only is this the best model we've ever tested," ARC Prize Foundation President Greg Kamradt said as part of the OpenAI announcement, "but it also represents a meaningful step change in frontier-model performance — not only in its ability to navigate and solve novel environments but also in how efficiently it learns to do so."
How much can we read into test results?
"When we launched ARC-AGI-3, "we made it clear that saturating the benchmark would not represent 'proof of achieving AGI,'" Kamradt wrote in a blog post. "Therefore, while we believe Astra represents meaningful progress towards generalization, we are not claiming that it is AGI."
Anka Reuel and Mike Hardy, doctoral researchers in AI at Stanford University, also raised concerns over OpenAI quietly revising several published metrics post-launch, including halving Astra's reported hallucination rate from 4.2% to 2.0% before reverting it.
ARC Prize Foundation President Greg Kamradt
In a 2025 paper published on pre-print server arXiv titled ‘Benchmarking is Broken -- Don't Let AI be its Own Judge’, experts including Princeton doctoral candidate Zerui Cheng and University of Luxembourg postdoctoral researcher Dr. Stella Wohnig warned that such practices create confusion and undermine trust, particularly when official technical documentation omits the basic methodology, making independent verification nearly impossible. Crucially, scores achieved through benchmaxxing reflect idealized performance ceilings under optimal lab conditions rather than practical, out-of-the-box reliability.
Highlights include a 98% score on FrontierMath Tier 4 (v2), which evaluates expert-level mathematical reasoning; a perfect 100% on ExploitBench, which is designed to test offensive cybersecurity capabilities; 95.9% on BenchCAD, which measures an AI's ability to reconstruct 3D objects using CAD code; 57.9% on Terminal-Bench 4.0, which assesses complex terminal-based system administration and software engineering; and 41.4% on AutomationBench, which tests whether agents can complete multistep business workflows across applications.
While some benchmark results have seen improvement since the previous generation, Astra underperformed in other tests. Artificial Analysis' testing showed that against GDPval-AA v2 — an economically focused benchmark that measures real-world workplace tasks across 44 occupations — Astra suffered a significant drop in its relative leaderboard ranking compared with GPT-5.6 Sol. Additional regressions were observed in benchmarks that test customer service support, scientific Python programming, and long-context reasoning across large documents.
AI systems today can do more than a simple back-and-forth text exchange, graduating to taking actions on our behalf. (Image credit: CFOTO via Getty Images)
On professional workplace evaluations like Agents' Last Exam, the new model reduced output token consumption by up to 65% compared with Opus 5, while on general intelligence evaluations, Astra achieved a roughly 10% output token reduction at max effort compared with Sol.
Does achieving AGI even matter?
With its latest model, OpenAI has placed a stronger emphasis on safeguards and controls. The move follows a series of high-profile incidents in which frontier AI models inadvertently hacked into third-party networks and computer systems as part of routine testing operations, after the models went beyond what humans expected they would do when encountering challenging tasks.
David Wood — chair of London Futurists, a non-profit that hosts discussion groups on emerging technology — argued that beyond individual test results, Astra highlights a fundamental shift in how humans will interact with AI as models move from answering questions to autonomously pursuing complex goals in digital environments.
Related stories"The impressive benchmark results matter, but what matters more is the combination of intelligence, autonomy, computer use, and cybersecurity capability," he said in an email to Live Science. "That combination also demands caution. The more useful these systems become, the greater the consequences when they misunderstand our intentions, are misused, or find ways around the safeguards we give them."
"The possibility that AI could progress from systems like Astra towards all-round superintelligence makes much greater human vigilance, awareness, and collaboration increasingly urgent," he said.
Help us improve Live Science Pro: We're always trying to make our content better. Leave us feedback about Pro here.
Hence then, the article about openai claims we ve entered the agi era has gpt 6 astra really demonstrated general intelligence was published today ( ) and is available on Live Science ( Middle East ) The editorial team at PressBee has edited and verified it, and it may have been modified, fully republished, or quoted. You can read and follow the updates of this news or article from its original source.
Read More Details
Finally We wish PressBee provided you with enough information of ( OpenAI claims we've entered the AGI era — has GPT-6 Astra really demonstrated general intelligence? )
Also on site :
- Witney Carson Claps Back at ‘DWTS’ Judge Carrie Ann Inaba, Joining Uproar Over Comments on Guillermo Rodriguez: ‘I Said What I Said’
- Disney Cruise Line Offers First Look Inside Disney Believe with New Frozen Dining, Encanto Show, and Themed Lounges
- Tom DeLonge Was in the Crowd When 43,000 Fans Serenaded Him With His Own Song: ‘He Was Definitely in His Feels’