More on the topic...
Generating detailed summary...
Failed to generate summary. Please try again.
Anthropic’s team set up nine copies of Claude Opus 4.6 as “Automated Alignment Researchers” (AARs). Each AAR got a sandbox, a shared forum, storage for code, a remote server for scoring, and a bit of background on training and inference. They faced a weak-to-strong supervision task: a weak model acts as teacher to fine-tune a stronger base model, and success is measured by how much of the “performance gap recovered” (PGR) the strong model achieves—0 means it matches the weak teacher, 1 means it hits its theoretical best. Researchers compared AARs to two humans who spent seven days on four known generalization methods and achieved a PGR of 0.23 using Qwen models.
Over five days and about 800 research hours, the AARs hit a PGR of 0.97, costing roughly $18,000 in token and training expenses. When the top two AAR methods moved to unseen tasks, one held up well on math (PGR 0.94) and less so on coding (0.47), while the second method did okay on math (0.75) but worsened coding performance. In a final test on Claude Sonnet 4 with production infrastructure, the leading AAR method showed no significant gain. Anthropic suggests this reflects the early trial’s limits and notes that AARs tend to exploit quirks of their specific models and data.
Running variants taught the team how to steer AARs for better results. Giving each AAR a different, even vague, starting prompt prevented them from converging on the same dead-ends and boosted progress. By contrast, imposing rigid workflows stifled creativity. When left to chart their own paths, AARs designed cheap initial experiments, then scaled up the promising ones. These findings guide future efforts to apply automated researchers across multiple domains and datasets.
Questions about this article
No questions yet.