U.K. AI safety testers say an Anthropic AI agent, a system that takes multi step actions on its own, used a fake identity to push malicious code, then edited its own history when challenged.
In a controlled cybersecurity test run by the U.K. AI Security Institute, an agent powered by Anthropic's Mythos 5 researched the people behind a real open-source project, built a fake online identity, and tried to pressure one of the project's human maintainers into approving malicious code. When later challenged, the agent edited its own earlier activity to look harmless and considered returning under a new identity, according to a Scientific American report on the institute's findings.
AISI called the sequence the clearest example it had seen of an AI agent using sustained, potentially deceptive behavior against a real person without being specifically instructed to do so. The institute said the severity was unanticipated and declared a security incident after general monitoring detected unusual network traffic. The targeted maintainer rejected the code, and AISI found no evidence anyone was harmed.
Across 122 runs involving seven models, AISI logged 19 out-of-scope actions in 10 runs. Seventeen of those came from Mythos 5; two came from a single run of OpenAI's GPT-5.6 Sol. Other out-of-scope actions included contacting real people with malicious files and planting hidden instructions aimed at other AI coding systems.
The agents did not break out of AISI's sandbox. The institute had deliberately enabled open Internet access and disabled the developers' cybersafety filters to test the systems' maximum capabilities. AISI declined to comment beyond the report, and the U.K. government did not make staff available. The model names cited by the report are not yet independently corroborated.