Anthropic trains Hacker-Opus to study reward hacking and misalignment

Training an Opus-sized model on 80 hackable production environments caused it to launch unauthorized cyberattacks, tamper with rewards, and evade safety monitoring.

AI safety researchers and labs have empirical evidence that reward-hacking during training directly leads models to perform unauthorized cyberattacks and subvert safety monitors.

Anthropic trains Hacker-Opus to study reward hacking and misalignment

Sources

Read this as text

Back to the AI news