A researcher in a bright yellow beanie working at a laptop covered in sticky notes at a kitchen table in warm evening light, editorial photography, shallow depth of field.
News

Anthropic's Automated Researchers Now Fix Alignment Failures Better Than Humans

Anthropic's fellows program published a paper showing automated AI systems reliably improved alignment performance on all 10 test benchmarks, sometimes proposing methods that beat experienced human researchers.

AnthropicAI safetyAutomated researchRecursive self-improvementAlignment

Anthropic published a paper on Friday with a quietly radical claim buried in its cost table: an AI system doing alignment research at roughly $4 per hour in API inference, against the $150 per hour the company pays its human researchers. The paper, Automated Researchers Can Reliably Mitigate Alignment Failures, is led by Chen Yueh-Han from Anthropic’s fellows program, and it is the most concrete look yet at what automated AI research actually looks like in practice.

The headline result: given 10 benchmarks targeting specific misaligned behaviours, automated systems improved performance on every single one without degrading the model’s overall capability. That second clause matters. Plenty of systems can pass alignment tests if you sand off general competence in the process. This one claims not to.

How the automated researcher works

The system mirrors the rhythm of a human research lab. Each automated researcher searches the existing alignment literature, proposes a training method, and applies it to the target model for a 30-minute training run. Then it measures. Methods that move the benchmark get kept; methods that don’t get discarded, and the loop repeats across several iterations.

The paper is candid about what that loop delivers. “The best AAR method beats what experienced humans propose, on average within six hours,” it reads — followed by the sharper twist: “Human guided research directions do not lead to stronger performance.” In other words, the humans weren’t just slower. Their research intuitions were worse.

The part people will argue about

There are two readings of this paper, and both deserve airtime.

The optimistic one: alignment post-training is becoming automatable. The paper’s own framing is that “automated alignment post-training could become practical in the near term.” If the task of making models behave can be handed off to cheap, tireless automated researchers, the safety work scales with the models instead of lagging behind them.

The uncomfortable one: this is a step toward recursive self-improvement, with the recursive part pointed squarely at the research labour itself. If models can improve their own alignment training, the same machinery can plausibly improve training practices more broadly. Anthropic has already said Claude writes most of its own code. Research direction is the next rung up, and this paper suggests the rung is load-bearing. What stands out here is the framing: Anthropic is publishing evidence that its own researchers’ judgement is being outperformed — by systems it built. That’s either reassuring or ominous depending on how much you trust the loop’s overseer.

The limitations the paper admits

Three big ones. The automated system is only as good as its benchmarks — if a benchmark doesn’t actually measure the behaviour you care about, the system optimises the proxy, not the goal. Someone still has to build and maintain those benchmarks, and that work is stubbornly human. The literature the automated researchers draw from is also curated by people, so the system inherits whatever blind spots live in that corpus.

None of these sink the approach. They just relocate the human problem upstream: from doing the research to deciding what the research should be aiming at.

For the broader context on why this is hard in the first place, our explainer on alignment covers the basics. And Anthropic isn’t alone here — Ornith AI generated headlines last week with models that write their own training curricula, and Sakana AI has been running its own recursive self-improvement lab for months. What’s new in the Anthropic paper is the safety focus: everyone else is automating capability. Anthropic is automating the part that keeps capabilities in check.

NZ angle: if automated alignment research becomes standard practice, the safety-audit work that regulators — including New Zealand’s approach under the AI strategy framework — assume requires human experts gets cheaper to verify. Independent auditing of AI systems could cost cents per benchmark instead of consultant day-rates. That’s a genuine shift in who can afford to check the models.

— CJ Murden, editor of Singularity.Kiwi. Former digital technologies teacher, author of AI-focused books. Writing with a New Zealand focus.

Sources: https://www.anthropic.com/research/automated-researchers-mitigate-alignment-failures, https://techcrunch.com/2026/08/28/an-anthropic-researcher-just-gave-us-a-peek-at-self-improving-ai/