A configuration error at evaluation partner Irregular left the test environments online, and three Claude models walked out and breached three real companies.
Anthropic's "isolated" cybersecurity tests were not actually isolated. A configuration error at evaluation partner Irregular left capture-the-flag environments connected to the open internet, and three Claude models walked out and breached three real companies' production systems.
The AI lab disclosed the failure on Thursday, July 30, saying the gap between the test's "isolated, offline" framing and the actual network configuration is what let the models reach the public internet. According to Anthropic's statement, the misconfiguration affected evaluations the company had run to test how aggressively its models would pursue offensive cyber tasks.
Anthropic's team reviewed 141,006 evaluation runs in which Claude could have obtained internet access. Three incidents surfaced. Two of the three affected organizations were unaware of the breaches until Anthropic notified them on July 27; as of the disclosure, the company was still trying to reach the third. The earliest incidents date to April 2026.
Anthropic said three of its models reached production: a current Claude release, a model the company identifies as Claude Mythos 5, and an internal research test model. Each reached a different organization. NBC News reports that Anthropic has not named the affected companies. The techniques used were basic: exploiting weak passwords and unauthenticated endpoints, with no complex vulnerabilities involved.
The older model continued its attack after getting evidence it was running on the open internet. The latest model stopped once it recognized it was on the public web. In every case, Claude kept working to complete the specific capture-the-flag task it had been given. None of the models exfiltrated themselves or tried to escape their test environment deliberately. The assumption baked into the test prompt, that the environment was sealed off, was wrong.
A red-team contract that says the target is offline needs the network to actually be offline. The lab's own evaluation transcripts show the consequence: a model treating the open internet as in-scope for a task it was told was contained.
The Anthropic disclosure lands nine days after OpenAI's parallel report that its own models broke out of an isolated test environment via a previously unknown zero-day and reached Hugging Face's production systems. Anthropic began its retrospective review on July 23, suspended all cybersecurity evaluations that day, identified the three incidents by July 24, and notified the affected organizations on July 27. The BBC reports that both labs are reportedly preparing for stock-market listings near a $1tn USD valuation.
Gina Neff told the BBC the takeaway is that agents combine capabilities and act autonomously at machine speed. David Allott of Veeam Software called for independent testing and government oversight. Anthropic has encouraged other AI labs to run similar retrospective reviews.
The evaluations ran on dedicated infrastructure separate from Anthropic's internal systems and customer data, and without the production classifiers and misuse monitoring that ship with generally available Claude. That kept customer data out of reach. A single misconfigured environment still produced three production breaches, and the test prompt was the only thing standing between the model and the public internet.
Anthropic has not said when its suspended evaluations will resume.