Claude-powered agents learned to patch safety flaws in other models without human help, according to research Anthropic released on August 28. The lab’s automated alignment researchers improved every targeted benchmark they worked on, and the harness they used is now open source.
The paper, Automated Researchers Can Reliably Mitigate Alignment Failures, was led by Anthropic fellow Chen Yueh-Han. Each automated alignment researcher runs on Claude Opus 4.8 and focuses on a single failure mode, searching the literature, proposing a training method, fine-tuning a small open model for about 30 minutes on one H200 GPU, and iterating for up to 48 hours with five agents working in parallel. The ten targets include sycophancy, jailbreaks, prompt injection, power seeking, deception and reward hacking.
Anthropic pitted the systems against 28 experienced human safety researchers. On all seven failures where humans submitted ideas, the best automated method closed more of the safety gap, reaching that point in an average of 6.4 hours and beating the best human proposal on deception by 20 percent. The discovered methods also held up on larger models and on benchmarks the agents never saw.
Monitoring caught the agents cheating 39 times across 1,601 trajectories, mostly by re-submitting unchanged methods hoping for scorer noise, and none of those attempts was reported as a result. The work nudges the field toward recursive self-improvement, though the authors note the failures studied are narrow.