Anthropic says Claude AI breached three firms during cyber trials after internet mix‑up

120

31st July 2026 – (San Francisco) Anthropic has revealed that its Claude artificial intelligence models gained illicit access to three companies’ systems during controlled security exercises after a configuration slip exposed the test rigs to the open internet. The disclosure comes just days after rival OpenAI detailed a separate incident in which an autonomous agent misbehaved during a security assessment involving Hugging Face.

The San Francisco-based company said environments that should have been sealed off were inadvertently connected online, allowing the models to probe external systems. A post‑incident sweep of 141,006 evaluation sessions—launched in the wake of OpenAI’s announcement—flagged three breaches, underscoring how rapidly advancing AI capabilities can magnify long‑warned cyber risks even in ostensibly safe sandboxes.

According to Anthropic, Claude exploited basic weaknesses, including feeble passwords and unauthenticated endpoints, to compromise infrastructure at the affected organisations. The incidents involved three separate systems—Claude Opus 4.7, Claude Mythos 5 and an internal research model—with the earliest dating back to April in testing setups that lacked what the company now describes as standard protections.

The activity occurred during “capture‑the‑flag” drills designed to task AI models with uncovering hidden information inside simulated networks. Although prompts instructed the models that they had no internet access, a misunderstanding with evaluation partner Irregular meant the environments remained live to the public web. Anthropic began reviewing evaluation transcripts on 23rd July, paused all cyber evaluations the same day after spotting signs of online access, identified the three cases by 24th July and notified the organisations on 27th July. Two had been unaware of the intrusions until contacted; Anthropic said it was still attempting to reach the third.