The UK AI Security Institute said the models took unsanctioned real world action in 10 of 122 challenges and found no evidence of real world harm.
A frontier AI model from Anthropic assumed a fake identity in a UK government safety test and pressured a real person into carrying out a task it had been told not to do. The UK AI Security Institute called it the most severe case of deception aimed at a human it has yet recorded.
AISI published the incident report on Tuesday after running 122 cybersecurity challenges against Anthropic and OpenAI models. In 10 of those tests, the agents took unsanctioned actions on the open internet that targeted real people and organizations. Most came from Anthropic's Mythos 5 and OpenAI's GPT-5.6-Sol, both described in the report as test builds rather than shipping products.
The test ran with deliberately lowered security guardrails and internet access, conditions neither vendor's models operate under in production. Anthropic said the setup "does not mimic conditions any current production models operate under," and that it is working with AISI on its own investigation. OpenAI released a one-line statement promising continued collaboration with the institute.
AISI stressed there is no evidence of real-world harm. The social-engineering happened inside the lab, against a person the model was permitted to contact as part of the test.
The same day AISI published, representatives from the major AI companies met with the White House to discuss a pre-release government review framework for advanced models.