#
language-models
→
automated-research
→
alignment-research
→
scalable-oversight
→
weak-to-strong-supervision
1 link tagged with all of: language-models + automated-research + alignment-research + scalable-oversight + weak-to-strong-supervision
Links
Nine copies of Claude Opus 4.6 were equipped with sandbox environments and tasked to autonomously develop weak-to-strong supervision methods, scoring their progress by “performance gap recovered” (PGR). The AARs reached a PGR of 0.97 versus a human baseline of 0.23, showed partial generalization to new tasks, but failed to replicate gains at production scale, underscoring both the promise and limits of automated alignment experiments.
- Nine Claude Opus 4.6 copies working autonomously hit a PGR of 0.97 on weak-to-strong supervision, crushing the human baseline of 0.23, in just 5 days and ~$18,000.
- The top method didn't generalize well: it worsened coding performance on unseen tasks and showed no significant gain when tested on Sonnet 4 with production infrastructure, suggesting it overfit to quirks of the original setup.
- Giving each AAR a different, even vague, starting prompt outperformed rigid, uniform workflows by avoiding convergence on the same dead-ends.
automated-research
weak-to-strong-supervision
scalable-oversight
language-models
alignment-research