1 link tagged with all of: ai-benchmarking + arc-agi + gpt-6 + agentic-intelligence + reasoning-efficiency
Links
OpenAI's GPT-6 Astra achieved near-perfect scores on ARC-AGI-3, a benchmark measuring agentic intelligence through novel puzzle environments, and matched human efficiency by solving 96% of levels with fewer actions than the median human participant.
- Astra scored 62.7% with standard evaluation but 99.9% when using OpenAI's provider-specific context management features, showing significant performance variation based on how the model manages information between requests.
- The model surpassed human action efficiency on 96% of levels, using 51.7% fewer actions per level on average—a shift from the long-held assumption that AI would require more exploration than humans to solve problems.
- Astra spontaneously developed compact algebraic notation and domain-specific languages to track game states, building custom symbolic world models for each environment without explicit instruction to do so.
ai-benchmarking
agentic-intelligence
gpt-6
arc-agi
reasoning-efficiency