
After OpenAI, Anthropic finds Claude breached three organizations during cyber tests
Less than two weeks after OpenAI disclosed that an experimental AI model breached Hugging Face during a cybersecurity evaluation, Anthropic has revealed that its own review uncovered three incidents in which Claude models gained unauthorized access to the production infrastructure of three organizations during similar testing.
Anthropic said it launched the review after OpenAI disclosed that one of its experimental models had escaped its evaluation environment.
“After reviewing 141,006 evaluation runs where Claude could have obtained internet access, we identified three incidents in which a model accessed the internet from within or while interacting with the evaluation environment of Irregu...