Anthropic says Claude Opus 5's layered defenses push prompt injection attacks — hidden commands smuggled into documents or web pages to hijack an AI — to roughly 0 percent success, in a vendor self report.
Anthropic says Claude Opus 5, the company's newest flagship model, is the hardest one yet to manipulate with text instructions. Prompt injection is a category of attack where someone slips hidden commands into an email, document, or web page to hijack an AI assistant's behavior.
Boris Cherny summarized the claim in a post on X and via Simon Willison's blog. Across the company's prompt-injection evaluations and red-team probes, Opus 5 is "very hard to prompt inject successfully." When defenders stack the model's own alignment with a prompt-injection probe and Claude Code's Auto Mode runtime policy, the attack-success rate drops to roughly 0 percent. The supporting detail sits on page 73 of the Opus 5 system card, a public safety report Anthropic published at launch.
The "roughly 0" figure describes the layered defense stack, not the model standing alone. The number is also a vendor self-report, not an independent benchmark. VentureBeat reported this week that three AI coding agents leaked secrets after a single prompt injection, and one of those vendors had flagged the vulnerability in its own system card. That is the gap between the lab's eval and the field track record: "least prompt injectable" is not "not prompt injectable," and prompt injection as a class of attack remains open.