Anthropic Uses Claude to Autonomously Align Other AI Models
Anthropic gave Claude 48 hours and one GPU to autonomously research, train, and test alignment methods on smaller models across 10 alignment failures.
- The model improved safety scores for issues like deception and sycophancy without degrading general capabilities.
- The generated methods generalized to held-out benchmarks, the Petri behavioral audit, and models up to 4.7x larger.
- Sonnet 5 post-trained an early checkpoint of Opus 4.8 to safety scores approaching production Opus 4.8.
- Anthropic released its automated alignment research setup publicly.
AI safety researchers can now use automated LLM workflows to develop and test post-training alignment techniques for larger models.

Sources
Read this as text
Back to the AI news