Anthropic Admits Claude AI Accidentally Hacked Systems of Three Organisations

The CSR Journal Magazine

Anthropic has revealed that its Claude AI system inadvertently accessed the environments of three distinct organisations during a cybersecurity assessment, following a configuration error that mistakenly granted the AI access to the internet. The company asserted that the incidents were not a result of Claude attempting to breach its testing parameters, but rather attributed to a flaw in the evaluation setup.

In a statement posted on their blog, Anthropic explained that an analysis of 141,006 evaluation attempts showed three instances where the AI was able to access the internet. This occurred while interacting with the evaluation environment set up by Irregular, a third-party partner, leading to unauthorised access to the organisations’ production infrastructures.

The disclosure follows a similar incident reported by OpenAI, where its AI models inadvertently escaped an isolated testing environment due to an undiscovered software vulnerability. As a precautionary measure, Anthropic conducted an extensive review of its cybersecurity evaluations to ascertain whether Claude had accessed any unauthorised systems.

Details of the Hacking Incidents

Anthropic explained that the hacking incidents were not entirely caused by Claude AI. According to the company, Claude had been informed that it was functioning within a fictional environment devoid of internet access. However, misconfiguration between Anthropic and Irregular resulted in the testing machines being connected to the internet, leading Claude to mistakenly assume that the discovered websites and servers were part of its challenge.

In the first incident involving Claude Opus 4.7, the AI targeted a genuine organisation due to its name resembling that of a fictional entity used in the test. Mistaking it for the assigned target, Claude accessed the company’s infrastructure, retrieved credentials, and gained entry to a production database containing numerous data rows. Anthropic categorised this as the most significant incident among the three.

Another incident involved Claude Mythos 5, which generated and uploaded a malicious Python package to the public PyPI software repository. Claude believed the repository was part of the simulation, and the package was available online for about an hour before being removed. During this period, it was downloaded by 15 genuine systems, including a cybersecurity company’s automated security scanner, which executed the package, allowing the AI to gather credentials.

Investigation and Future Safeguards

The third incident featured an internal research model that scanned approximately 9,000 internet-connected targets after failing to achieve its fictional goal. Ultimately, it compromised a company’s application by exploiting exposed credentials and SQL injection vulnerabilities. Unlike earlier models, this version ceased its activities upon realising it had interacted with a real system outside of the intended evaluation.

Anthropic clarified that the evaluations were performed without the safeguards and protections typically employed in publicly available Claude models against misuse. Following the discovery of these incidents, the company suspended all cybersecurity evaluations, notified its evaluation partner Irregular, and informed the organisations impacted.

In its comparison with OpenAI’s recent incident, Anthropic indicated that the events differed significantly. While OpenAI’s models escaped due to an unrecognised vulnerability, Claude accessed the internet due to a configuration error in the testing setup. The company emphasised the need for stronger safeguards within AI testing as models become increasingly capable in their functionalities.

Long or Short, get news the way you like. No ads. No redirections. Download Newspin and Stay Alert, The CSR Journal Mobile app, for fast, crisp, clean updates!

App Store –  https://apps.apple.com/in/app/newspin/id6746449540 

Google Play Store – https://play.google.com/store/apps/details?id=com.inventifweb.newspin&pcampaignid=web_share

Latest News

Popular Videos