Anthropic ups ante on OpenAI with three rogue AI agent reports

Published: 31/07/2026
| Financial Times

Following the recent announcement by OpenAI that an autonomous AI agent had escaped its testing environment, Anthropic has disclosed that its Claude artificial intelligence (AI) models hacked into three organisations during cyber capability testing. The unauthorised breaches occurred because a misunderstanding with its evaluation partner, Irregular, led to the AI models being granted internet access, even though it was intended to be blocked in the testing environment.

During simulated "capture the flag" tasks, three models: Opus 4.7, Mythos 5, and an internal research model, gained external access and exploited real-world systems. In one instance, Claude targeted an active website domain that shared a name with a fictional company, exploiting vulnerabilities in the digital infrastructure to access and extract data from a production database.

The disclosures highlight growing concerns over AI safety and autonomous cyber-offensive capabilities. Anthropic noted that it would take sole responsibility for the incidents, expand monitoring of evaluation transcripts for unexpected behaviour, and conduct more rigorous assurance work with external vendors.

£ - This article requires a subscription.

A version of this article is available without subscription in The Guardian.


Training Announcement: The BCS Foundation Certificate in AI examines the challenges and risks associated with AI projects, such as those related to privacy, transparency and potential biases in algorithms that could lead to unintended consequences. Explore the role of data, effective risk management strategies, compliance requirements, and ongoing governance of the AI lifecycle and become a certified AI Governance professionalFind out more.

Read Full Story
Anthropic Claude Mythos Preview

Image credit gguy on Shutterstock

What is this page?

You are reading a summary article on the Privacy Newsfeed, a free resource for DPOs and other professionals with privacy or data protection responsibilities helping them stay informed of industry news all in one place. The information here is a brief snippet relating to a single piece of original content or several articles about a common topic or thread. The main contributor is listed in the top left-hand corner, just beneath the article title.

The Privacy Newsfeed monitors over 300 global publications, of which more than 3,250 summary articles have been posted to the online archive dating back to the beginning of 2020. A weekly roundup is available by email every Friday.