More on the topic…
OpenAI released GPT-6 Astra today, claiming it marks the beginning of the AGI era. The real story for businesses is simpler than the AGI framing: Astra can use computers the way humans do—navigating browsers, filling forms, manipulating spreadsheets, writing code—without requiring custom API integrations for each application. Instead of telling users how to complete tasks, it actually completes them. Astra rolls out Thursday to enterprise customers through OpenAI's Daybreak program, then to ChatGPT Plus/Pro/Business users and via AWS Bedrock and Azure. The shift matters because companies have spent years building connectors and plugins to link AI models to their tools. If an AI can just use the existing human interface—keyboard, mouse, screen—that integration work disappears. On an offline subset of OSWorld 2.0, Astra scored 72.6% in roughly 40 minutes per task, compared to GPT-5.6 Sol's 65.7% at 75 minutes.
The training behind Astra was OpenAI's largest yet, using over 100,000 DBUs on the Stargate infrastructure and having previous models supervise the training of the next generation. According to researcher Aidan Clark, the jump from Sol to Astra represents a larger capability increase than Sol's jump over previous models. The benchmark results are substantial: 97.6% on FrontierMath Tier 4 v2, 74.1% on DeepSWE v1.1, 96% on GPQA Diamond, and 100% on ExploitBench. But there's a complication with the 98.6% score on ARC-AGI-3, which is supposed to measure whether AI systems can generalize to unfamiliar problems rather than just reproduce trained capabilities. The score looks impressive until you compare it to NVIDIA's recent 100% result on the same benchmark—which came from Claude Opus 5 (which scored only 30% baseline) wrapped in elaborate agent architecture with persistent memory and feedback loops.
This gap reveals a growing problem in how the industry measures AI progress. NVIDIA's result showed that long-horizon capability can emerge from the complete agent system rather than the foundation model alone, making it unclear whether you're measuring the underlying model's actual generalization or the sophistication of the harness around it. The AI community is already divided on what this means. Some argue that adding elaborate systems misrepresents the model's true capabilities, while others counter that ARC-AGI-3's restrictions on retaining context make it unrealistic compared to how production agents actually work. The practical question for enterprises is more concrete: does Astra actually reduce the integration work companies do, and by how much? The AGI debate is secondary to whether it actually saves money and time on real business workflows.
Questions about this article
No questions yet.