Skip to main content
Back to News Hub
🟢TechCrunch AI
July 21, 2026
ChatGPT

OpenAI says Hugging Face was breached by its pre-release models

Overview

OpenAI has come forward to claim responsibility for the Hugging Face breach, saying it was the result of internal testing gone awry. OpenAI admitted Tuesday that one of its AI models breached the systems of Hugging Face, the unaffiliated AI hosting platform, during an internal cybersecurity test that went awry. The models reportedly escaped their isolated testing environment and reached Hugging Face's systems from there.

Key Takeaways

  • Hugging Face initially attributed the breach to an "external AI agent."

    In a blog post published Tuesday afternoon , OpenAI detailed the steps that led the models to compromise the service.

  • Benchmarks like ExploitGym are commonly used in model training to refine specific skills, but this is the first known incident in which that testing resulted in an actual cyberattack.

    In this case, the model in question should not have even had internet access, outside of a specific tool that enabled models to install software packages they might need to complete their task.

  • Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation."

    Ultimately, the models found vulnerabilities in Hugging Face's infrastructure that allowed them to "obtain test solutions directly from Hugging Face's production database," effectively providing the answers to the benchmark.

  • It's unclear whether OpenAI will face any legal consequences as a result of the breach, although it's likely that the models' actions violated the Computer Fraud and Abuse Act.

    Nevertheless, the result is an unusually vivid illustration of the power and dangers of frontier AI models operating on long time horizons.

  • He can be reached at russell.brandom@techcrunch.com or on Signal at .

Hugging Face initially attributed the breach to an "external AI agent." In a blog post published Tuesday afternoon , OpenAI detailed the steps that led the models to compromise the service. "After investigating, we now know that this particular incident was driven by a combination of OpenAI models - including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes - while being internally tested on a benchmark⁠ of cyber capabilities," the post reads.

In particular, the breach appears to have focused on ExploitGym , a publicly hosted benchmark measuring models' ability to execute attacks based on existing vulnerabilities. Benchmarks like ExploitGym are commonly used in model training to refine specific skills, but this is the first known incident in which that testing resulted in an actual cyberattack. In this case, the model in question should not have even had internet access, outside of a specific tool that enabled models to install software packages they might need to complete their task.

Instead, the model was able to find an undisclosed vulnerability in the package-installer program, which it used to access the broader internet at will. "The models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal," OpenAI's post reads. "After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym.

For more details please read the original article at TechCrunch AI.

Continue Learning

Comments

Comments appear only after moderation. Your email identifies your submission to the moderator and is never displayed here.

No approved comments yet.

Originally published by TechCrunch AI
Read the original