More on the topic…
Anthropic is documenting a shift where AI systems are taking over more of AI development itself. They're showing that AI can now handle coding tasks that used to require humans—writing entire files, running code autonomously, and delegating work to other agents. The company argues this trend could eventually lead to recursive self-improvement, where an AI system designs and trains its own successor. They're not claiming this is inevitable, but they're presenting evidence it's happening faster than most people realize. The risks are real: if AI systems can fully build themselves, controlling and monitoring them becomes exponentially harder.
The data backs up the acceleration claim. Claude went from completing 4-minute tasks in March 2024 to 12-hour tasks by mid-2026—roughly doubling capability every four months. On real benchmarks, the jump is even starker. SWE-bench, which tests actual software engineering on real codebases and bugs, went from single-digit performance to saturation in two years. CORE-Bench, measuring whether AI can reproduce published research, climbed from 20% success to saturation in fifteen months. These aren't theoretical metrics—they measure what AI can actually do in the wild.
Inside Anthropic itself, the productivity shift is concrete. Over 80% of code merged into their codebase in May 2026 was written by Claude, up from single digits before February 2025. Engineers are shipping 8x as much code per quarter compared to 2024, though Anthropic acknowledges this overstates true productivity gains since quantity doesn't equal quality. A March 2026 survey of their research teams found the median person felt 4x more productive with AI assistance. The gap that remains is judgment—Claude can execute well-specified tasks and solve underspecified problems, but it still struggles with choosing which problems matter in the first place. That's the remaining barrier to full self-improvement.
Questions about this article
No questions yet.