1 link tagged with all of: automated-research + weak-to-strong-supervision + scalable-oversight + language-models
Click any tag below to further narrow down your results
Links
Nine copies of Claude Opus 4.6 were equipped with sandbox environments and tasked to autonomously develop weak-to-strong supervision methods, scoring their progress by “performance gap recovered” (PGR). The AARs reached a PGR of 0.97 versus a human baseline of 0.23, showed partial generalization to new tasks, but failed to replicate gains at production scale, underscoring both the promise and limits of automated alignment experiments.
- Nine Claude Opus 4.6 copies working autonomously hit a PGR of 0.97 on weak-to-strong supervision, crushing the human baseline of 0.23, in just 5 days and ~$18,000.
- The top method didn't generalize well: it worsened coding performance on unseen tasks and showed no significant gain when tested on Sonnet 4 with production infrastructure, suggesting it overfit to quirks of the original setup.
- Giving each AAR a different, even vague, starting prompt outperformed rigid, uniform workflows by avoiding convergence on the same dead-ends.
automated-research
weak-to-strong-supervision
scalable-oversight
language-models
+ alignment-research