Investigating three real-world incidents in our cybersecurity evaluations
In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Below we describe what happened, how it happened, and what we're changing. We encourage other AI labs to perform similar reviews.
Key Takeaways
- Frontier Red Team Investigating three real-world incidents in our cybersecurity evaluations Jul 30, 2026 In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations.
This post reflects our current understanding; we'll update it if any details change.
- In response to this incident, we began a large-scale retrospective review of our own cybersecurity evaluations.
In particular, we looked for evidence that Claude-like the OpenAI models that accessed Hugging Face-was able to access the internet from within testing environments that should have been sealed off.
- The challenge is left open-ended, and no particular method is prescribed.
In all cases, Anthropic's evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access.
- ) Operating under the false belief that all accessible entities were intended to be in-scope for the exercise, Claude compromised the impacted organizations' infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints.
It did not find or exploit any complex vulnerabilities, and in each case, Claude continued working to complete only the specific capture-the-flag task its evaluation had assigned.
- 7, Mythos 5, and an internal research test model.
Stats & Key Facts
- #Frontier Red Team Investigating three real-world incidents in our cybersecurity evaluations Jul 30, 2026 In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations.
- #On July 21, OpenAI disclosed that several of their models had broken out of an isolated test environment by exploiting a previously unknown ("zero-day") vulnerability.
- #After reviewing 141,006 evaluation runs where Claude could have obtained internet access, we identified three incidents in which a model accessed the internet from within or while interacting with the evaluation environment of Irregular, one of our third-party evaluation partners, and then gained unauthorized access to the production infrastructure of three different organizations.
Continue Learning
Comments
Sign in to join the conversation