AI Safety Regulations in the U.S. Could Give Hackers an Edge
On 11 July, Hugging Face was subjected to an intense cyberattack from a then-unknown actor. The speed and coordination of the attack on the company that hosts and supports popular AI developer resources led Hugging Face's security team to conclude it was the work of an AI agent. Realizing this, the team tried to use "frontier models behind commercial APIs"-presumably from Anthropic and OpenAI, although only Anthropic was named in the second of the company's two posts about the security incident-to analyze the onslaught.
Key Takeaways
- These models refused to help due to safety guardrails the AI labs have implemented to make their models harder to use for cyberattacks.
- " Massive AI Cyberattack on Hugging Face The scale of the OpenAI model's attack on Hugging Face was massive.
Across five days, it executed over 17,500 individual actions , such as privilege escalation and code execution.
- The model was ultimately successful in extracting five dataset files, though it's not clear if the data helped it achieve its goal.
OpenAI and Hugging Face did not respond to requests for comment.
- On 30 July, Anthropic disclosed three instances where a model executed an attack as part of an evaluation.
In one case, Claude uploaded malware to PyPI , the official Python software repository.
- The Scale AI team quantified the problem in a paper published at ICLR 2026 , which found that, depending on the task, nearly 44 percent of defensive requests were refused.
Stats & Key Facts
- #These models refused to help due to safety guardrails the AI labs have imp On 11 July, Hugging Face was subjected to an intense cyberattack from a then-unknown actor.
- #On 21 July, OpenAI announced the attacker was an OpenAI model undergoing testing in a sandboxed environment.
- #At its peak, the model performed more than 300 actions per hour.
- #On 30 July, Anthropic disclosed three instances where a model executed an attack as part of an evaluation.

On 21 July, OpenAI announced the attacker was an OpenAI model undergoing testing in a sandboxed environment. It escaped its internal sandbox, established a foothold in a third-party server, and then assailed Hugging Face. In other words, frontier models-those that score highest in AI performance benchmarks-had refused to assist Hugging Face's security team in analyzing the attack, yet a prospective frontier model in testing had executed it in the first place.
"I would argue that asymmetry is the paramount problem of our time," says Alex Levinson , executive director of the National Collegiate Cyber Defense Competition and coauthor of a paper on defensive refusal bias . "We want the world to exist in a state of security, but we're not going to get there by guardrailing away model capability. " Massive AI Cyberattack on Hugging Face The scale of the OpenAI model's attack on Hugging Face was massive.
Across five days, it executed over 17,500 individual actions , such as privilege escalation and code execution. At its peak, the model performed more than 300 actions per hour. While the attack resulted in little damage to Hugging Face's infrastructure, the model was able to steal credentials, gain admin access, and extract some data.
For more details please read the original article at IEEE Spectrum AI.
Continue Learning
Comments
Sign in to join the conversation