
Anthropic makes changes to stop AI agents running amok again
Learning from the OpenAI-Hugging Face fiasco, as well as from recent revelations about its own model, Anthropic is revamping its security and alignment practices.
The company has established controls that flag when a model attempts to break out of a sandbox or successfully accesses the live internet, cordoned off its highest-risk test environments, and proposed a set of safety standards for its external testing partners, such as giving AI agents explicit instructions like “you should not access the internet.”
Anthropic conceded that three recent security incidents involving Claude reflect a “failure of operational security,” and also reveal issues with model reasoning capabilities and “reckl...