OpenAI on Thursday released GPT-6 Astra, its newest flagship model, and President Greg Brockman declared it marks the arrival of artificial general intelligence. "Welcome to the AGI era," Brockman said during a press briefing, though he left it to observers to decide whether the model meets their own definition of AGI. The company calls Astra "the world's most intelligent and aligned model," and benchmark results show substantial gains over earlier versions, particularly in tasks outside traditional coding.

Astra scored 98.6% on the ARC-AGI-3 benchmark, a test of general reasoning, and achieved 97.6% on FrontierMath Tier 4's 41 private problems. On BenchCAD's Vision2Code subset, which asks models to reconstruct CAD programs from rendered views, Astra hit 95.9% compared to 84.3% for Anthropic's Fable 5.1 and 83.3% for OpenAI's own Sol. The model reached 64.6% on Terminal-Bench Science, a test of command-line research tasks across five scientific fields, versus 52.6% for Fable 5.1. On the DeepSWE v1.1 agentic coding benchmark, Astra scored 74.1%, up from 70.8% for Sol, though Meta's Muse Spark 1.3 reported 75.4% at its maximum reasoning setting. Astra will cost $10 per million input tokens and $50 per million output tokens through the API, matching Fable 5.1's pricing but sitting 2.5 times higher than Sol's promotional rate and far above competitors like Google's Gemini 3.8 Flash at $0.75/$3.75.

According to Aidan Clark from OpenAI, Astra represents the company's largest training run to date, marking the first time the firm pre-trained on more than 100,000 GPUs at its Stargate site in Texas. OpenAI also disclosed that earlier models played a significant role in supervising Astra's training process, a first for the company's model releases. On alignment, OpenAI Chief Scientist Jakub Pachocki stated that "progress in intelligence does not guarantee progress in alignment," adding that the company "will withhold scaling until we can regain enough confidence" in monitoring future models. The company reported Astra's written reasoning was harder to monitor than Sol's in tests designed to elicit monitoring evasion, though it went outside authorized targets in 0% of impossible-task scenarios compared to 48.2% for Sol without production safeguards.

The report reveals that Astra has crossed the Critical cybersecurity threshold in OpenAI's Preparedness Framework, developing exploits for hardened browsers and operating systems in company tests. It discovered two previously unknown vulnerabilities while being evaluated against recent V8 bugs, which OpenAI is disclosing to maintainers. Brockman emphasized that "the price per task is what matters" when justifying Astra's higher token costs, arguing that the model uses fewer tokens on several evaluations and in partner tests, though launch data is too sparse to confirm whether savings offset the premium. For developers, Astra introduces experimental features including the ability to keep notes across context windows and search earlier messages, plus the capacity to ask users questions without halting work that doesn't depend on the answer. The model's rollout begins with enterprise customers in OpenAI's Daybreak program before expanding to Plus, Pro, Business, and Enterprise users in coming days, though OpenAI's Mia Glaese warned that users may experience slowdowns, pauses, or blocks while performing cybersecurity work at launch. The timing of Brockman's AGI declaration, paired with Astra's demonstrable capability gains, shifts the conversation from whether such systems will arrive to how organizations should prepare for models that can complete a researcher's week of work in hours. For enterprises weighing adoption, the trade-off between raw capability and the opacity that comes with it will define deployment strategies far more than any benchmark score.