Skip to main content
Back to News Hub
🧠Anthropic
July 30, 2026
Funding & Investment

Investigating three real-world incidents in our cybersecurity evaluations

Overview

In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Below we describe what happened, how it happened, and what we're changing. We encourage other AI labs to perform similar reviews.

Key Takeaways

  • Frontier Red Team Investigating three real-world incidents in our cybersecurity evaluations Jul 30, 2026 In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations.

    This post reflects our current understanding; we'll update it if any details change.

  • In response to this incident, we began a large-scale retrospective review of our own cybersecurity evaluations.

    In particular, we looked for evidence that Claude-like the OpenAI models that accessed Hugging Face-was able to access the internet from within testing environments that should have been sealed off.

  • The challenge is left open-ended, and no particular method is prescribed.

    In all cases, Anthropic's evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access.

  • (Cybersecurity evaluation ranges commonly include realistic details in order to accurately assess what models are capable of in real settings; a realistic-looking target would not itself be clear evidence to a model that the target is not part of a simulation.)

    Operating under the false belief that all accessible entities were intended to be in-scope for the exercise, Claude compromised the impacted organizations' infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints.

  • The incidents involved three different Claude models: Opus 4.7, Mythos 5, and an internal research test model.

Stats & Key Facts

  • #Frontier Red Team Investigating three real-world incidents in our cybersecurity evaluations Jul 30, 2026 In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations.
  • #On July 21, OpenAI disclosed that several of their models had broken out of an isolated test environment by exploiting a previously unknown ("zero-day") vulnerability.
  • #After reviewing 141,006 evaluation runs where Claude could have obtained internet access, we identified three incidents in which a model accessed the internet from within or while interacting with the evaluation environment of Irregular, one of our third-party evaluation partners, and then gained unauthorized access to the production infrastructure of three different organizations.

Continue Learning

Comments

Comments appear only after moderation. Your email identifies your submission to the moderator and is never displayed here.

No approved comments yet.

Originally published by Anthropic
Read the original