OpenAI says Hugging Face was breached by its own pre-release models
OpenAI has come forward to claim responsibility for the Hugging Face breach, saying it was the result of internal testing gone awry. OpenAI admitted Tuesday that one of its AI models breached the systems of Hugging Face, the unaffiliated AI hosting platform, during an internal cybersecurity test that went awry. The models reportedly escaped their isolated testing environment and reached Hugging Face's systems from there.
Key Takeaways
- Hugging Face initially attributed the breach to an "external AI agent.
" In a blog post published Tuesday afternoon , OpenAI detailed the steps that led the models to compromise the service.
- In particular, the breach appears to have focused on ExploitGym , a publicly hosted benchmark measuring models' ability to execute attacks based on existing vulnerabilities.
Benchmarks like ExploitGym are commonly used in model training to refine specific skills, but this is the first known incident in which that testing resulted in an actual cyberattack.
- "After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym.
Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation.
- The company also said it would implement new controls on both model testing and the related infrastructure, meant to prevent similar incidents in the future.
It's unclear whether OpenAI will face any legal consequences as a result of the breach, although it's likely that the models' actions violated the Computer Fraud and Abuse Act.
- He previously worked at The Verge and Rest of World, and has written for Wired, The Awl and MIT's Technology Review.
Hugging Face initially attributed the breach to an "external AI agent. " In a blog post published Tuesday afternoon , OpenAI detailed the steps that led the models to compromise the service. "After investigating, we now know that this particular incident was driven by a combination of OpenAI models - including GPT‑5.
6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes - while being internally tested on a benchmark of cyber capabilities," the post reads. In particular, the breach appears to have focused on ExploitGym , a publicly hosted benchmark measuring models' ability to execute attacks based on existing vulnerabilities. Benchmarks like ExploitGym are commonly used in model training to refine specific skills, but this is the first known incident in which that testing resulted in an actual cyberattack.
In this case, the model in question should not have even had internet access, outside of a specific tool that enabled models to install software packages they might need to complete their task. Instead, the model was able to find an undisclosed vulnerability in the package-installer program, which it used to access the broader internet at will. "The models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal," OpenAI's post reads.
For more details please read the original article at TechCrunch AI.
Continue Learning
Comments
Sign in to join the conversation