Anthropic finds Claude accessed real systems in three evaluation incidents

Anthropic reviewed its cybersecurity evaluations and found three incidents where a Claude model gained unauthorized access to real systems of three organizations.

Organizations running AI evaluations must now account for the risk that a model under test can escape its environment and access real external systems.

Sources

Read this as text

Back to the AI news