
Anthropic Discovers AI Agents Given Conflicting Instructions Soon Tried to Sabotage Each Other
When Anthropic instructed three agents to migrate a Python backend, but telling each agent to perform the migration in a different language, "We consistently saw a multiagent turf war," they wrote Thursday:
All of the models we tested quickly assumed that others were purposefully impeding their work, and began to sabotage others while protecting their own contributions. In fact, they sabotaged others with increasingly aggressive, self-replicating malware. This included disabling the Unix accounts of the other agents, writing automated scripts that found and killed competing processes on a loop, and deploying malicious code that was disguised as belonging to another agent.
In many runs, on...