Back to News Hub
📐SiliconANGLE AI
July 31, 2026
Funding & Investment

Anthropic discloses that Claude hacked three organizations during internal tests

Overview

Three of Anthropic PBC's large language models carried out successful cyberattacks during routine internal tests. The company detailed the breaches on Thursday. A few days earlier, rival OpenAI Group PBC disclosed a similar incident.

Key Takeaways

  • SiliconANGLE UPDATED 17:21 EDT / JULY 31 2026 AI Anthropic discloses that Claude hacked three organizations during internal tests by Maria Deutscher Three of Anthropic PBC's large language models carried out successful cyberattacks during routine internal tests.

    Two of the company's LLMs escaped from an isolated sandbox that was being used to evaluate their cybersecurity capabilities.

  • Claude is tasked with finding a way of stealing data from the simulated organization's systems.

    Anthropic developed the test environments in collaboration with Irregular, an AI security startup.

  • 7 compromised a production database with several hundred rows of information.

    Additionally, it obtained access credentials for several applications and infrastructure assets.

  • At one point, the model discovered that the application wasn't a part of its security evaluation sandbox and stopped the cyberattack.

    Anthropic is partnering with a nonprofit AI safety lab called METR to carry out a more detailed investigation of the breaches.

  • com/aws-marketplace/ About SiliconANGLE Media SiliconANGLE Media is a recognized leader in digital media innovation, uniting breakthrough technology, strategic insights and real-time audience engagement.

Stats & Key Facts

  • #SiliconANGLE UPDATED 17:21 EDT / JULY 31 2026 AI Anthropic discloses that Claude hacked three organizations during internal tests by Maria Deutscher Three of Anthropic PBC's large language models carried out successful cyberattacks during routine internal tests.
Anthropic discloses that Claude hacked three organizations during internal tests

Two of the company's LLMs escaped from an isolated sandbox that was being used to evaluate their cybersecurity capabilities. They subsequently hacked Hugging Face, a popular platform for hosting open-source AI projects. OpenAI's disclosure prompted Anthropic to check logs from its own model security evaluations.

That review is what led to discovery of the cyberattacks disclosed on Thursday. According to Anthropic, its engineers identified three breaches carried out by three different Claude models. All three cyberattacks occurred during so-called capture the flag evaluations.

During such tests, Anthropic installs a Claude model in a sandbox that simulates the infrastructure of an external company. Claude is tasked with finding a way of stealing data from the simulated organization's systems. Anthropic developed the test environments in collaboration with Irregular, an AI security startup.

For more details please read the original article at SiliconANGLE AI.

Continue Learning

Originally published by SiliconANGLE AI
Read the original

Comments

Sign in to join the conversation