
Hugging Face breach shows why incident response needs a multi-model AI strategy
The recent breach of Hugging Face’s platform was the latest in a string of AI-assisted intrusions to come to light in recent weeks, showing that attackers can now use LLMs to automate entire attack chains. But it also exposed a limitation for defenders trying to use AI to respond at similar machine speed: Increasingly conservative safety controls on frontier models can block attempts to analyze malicious payloads and other intrusion artifacts that are needed to investigate incidents.
The incident was caused by an internal OpenAI test of advanced model cyber capabilities that went wrong. GPT-5.6 Sol and a more capable pre-release model, operating with reduced cyber refusals and without the pr...