OpenAI's AGI Claim for GPT-6 Astra Faces Skepticism from Researchers

OpenAI says its GPT-6 Astra model marks the beginning of the artificial general intelligence era. Independent researchers and benchmark creators argue that the model's record scores rely heavily on prompt optimization and that testing is not transparent. The company also reportedly delayed a planned GPT-6.1 Astra release because of safety concerns.
OpenAI framed GPT-6 Astra as a milestone on the path to AGI, pointing to leading results in coding, computer operation, mathematics, and science. It showed the model producing CAD code, laying out circuit boards, converting 3D models into playable environments, and launching hosted web apps from instructions. On ARC-AGI-3, which tests rapid skill acquisition in unfamiliar abstract settings, Astra reportedly reached 99.9%.
Outside evaluators challenge that narrative. They say the headline figures depend heavily on crafted prompts and that the evaluation process is not open enough. OpenAI also reportedly postponed a planned GPT-6.1 Astra launch because of safety worries, complicating claims that broad human-like intelligence has arrived.
If Astra's reported abilities hold, teams in software, design, legal work, and research could gain faster prototyping and task automation. Workers in those areas may see roles shift, while educators and auditors could struggle to distinguish genuine capability from benchmark-specific performance. If the skepticism proves correct, public confidence in AGI announcements may weaken and firms may delay deployments. The broader effect could be a stronger demand for independent, reproducible testing before major AI claims are widely accepted.