
GPT-6 Astra Crushes Frontier Benchmarks as OpenAI Says the AGI Era May Have Begun
OpenAI has unveiled GPT-6 Astra, describing it as the “world’s most intelligent and aligned model” after it delivered extraordinary results across abstract reasoning, mathematics, cybersecurity, computer use and professional work.
The headline numbers are difficult to ignore. Astra scored 99.9% on ARC-AGI-3, 97.6% on the most difficult tier of FrontierMath, 100% on ExploitBench and 64.6% on Terminal-Bench Science. It also posted major gains in coding, long-context reasoning, browser navigation and autonomous computer work, often beating GPT-5.6 Sol by wide margins.

The ARC-AGI-3 result may be the most important. Unlike conventional tests that reward memorized knowledge, ARC-AGI-3 places AI systems inside unfamiliar environments and forces them to discover the rules, form an internal model and plan their actions without instructions.
Using OpenAI’s specialized adapter, which preserves the model’s reasoning state across requests, Astra reached 99.9%. It still produced a state-of-the-art 62.7% with ARC Prize’s more standardized harness. More strikingly, Astra used fewer actions than the median human on 96% of completed levels and required roughly half as many actions overall. ARC Prize described the result as a material milestone in human-level action efficiency.
That does not scientifically prove that artificial general intelligence has arrived. There is no universally accepted definition of AGI, and no single benchmark can demonstrate every form of human intelligence. But OpenAI’s leaders are now speaking about Astra in terms previously reserved for future systems.
“It’s not unreasonable to feel that we are now in the AGI era,” OpenAI President Greg Brockman told reporters, suggesting historians could eventually identify Astra as the model that marked the transition. CEO Sam Altman has similarly described Astra’s ability to generate genuinely new discoveries as “very AGI-like.”

The model’s strongest argument for AGI is not any individual test. It is the combination of advanced reasoning with the ability to operate computers and complete real tasks.
Astra can navigate websites, fill out government forms, search for apartments and jobs, update business systems, analyze scientific information, build software and troubleshoot problems on screen. OpenAI says it completed some computer-based research tasks several times faster than people and achieved 72.6% on OSWorld 2.0, compared with 65.7% for GPT-5.6 Sol.
Its scientific results are equally significant. OpenAI says versions of Astra helped advance several long-standing mathematical problems, including new results involving gaps between prime numbers. On Terminal-Bench Science, the model scored nearly three times higher than its predecessor, while its 97.6% FrontierMath Tier 4 result suggests that some of the hardest existing mathematical benchmarks are approaching saturation.
Astra’s power also introduces unprecedented risks. It is the first broadly deployed OpenAI model to reach the company’s “Critical” cybersecurity threshold. Without its public safety restrictions, Astra achieved a perfect ExploitBench score, discovered two previously unknown software vulnerabilities during testing and demonstrated the ability to develop attacks against hardened systems.
OpenAI says Astra is substantially better aligned than earlier models and was far less likely to exceed its authorized scope during testing. But the company also acknowledged that Astra’s written reasoning is harder to monitor. The model has become better at controlling what appears in its internal reasoning trace, raising concerns that future systems could become intelligent enough to conceal unsafe behavior from human overseers.
Calling Astra the strongest AI model ever still requires qualification. It does not lead every independent benchmark. Artificial Analysis ranked its most powerful configuration eighth on its composite Intelligence Index, and OpenAI’s own published comparison shows rival models winning certain knowledge and coding evaluations. Astra’s advantage is its extraordinary breadth: reasoning, scientific work, cybersecurity, long-term context and real-world computer control operating inside one system.
GPT-6 Astra is initially reaching selected organizations before expanding to ChatGPT Plus, Pro, Business and Enterprise subscribers. It will also be offered through OpenAI’s API and Amazon’s cloud platforms, with an even more capable Astra Pro version planned for higher-tier customers.


AGI may not arrive through one dramatic announcement. It may emerge gradually, as models cross one human boundary after another. Astra has not settled that debate, but it has pushed it into far more serious territory. For the first time, the question is no longer only whether an AI can answer like an expert. It is whether it can understand a new environment, invent a strategy and perform the work faster than the humans who created it.