Adversaries targeted leading AI models' step by step reasoning through normal looking API traffic, scaling to 16,000 requests from more than 4,000 accounts in two days — with a related cluster spanning more than 15,000 more accounts before
A new attack on frontier AI reasoning went public this month, and OpenAI is the first frontier developer to put a number on it.
In a blog post and an accompanying research preprint, OpenAI says it disrupted a coordinated campaign in which adversaries tried to extract the internal reasoning of its models, the step-by-step working memory a model produces before it answers, by sending normal-looking requests through its public API. The company says no encryption was broken, no database was compromised, and no stored user conversations were touched.
The campaign began on July 1 at low volume, then spiked between July 24 and 25 with roughly 16,000 requests carrying an extraction pattern from more than 4,000 accounts, according to OpenAI's account. A related cluster of activity spread across more than 15,000 accounts before the company says it was fully disrupted by July 28.
The technique is what security researchers call a reasoning replay attack. Modern reasoning models encrypt the intermediate traces they generate so that downstream users cannot read them directly. OpenAI says the operators it observed were able to reuse those encrypted blocks across conversations, sessions, and accounts without breaking the encryption itself. Once the encrypted reasoning could be replayed at will, the attackers did not need to crack the strongest model in the family. Instead, they fed the traces to a weaker model inside the same provider ecosystem and used it to reconstruct the reasoning in readable form. The preprint describes interchangeable encrypted blocks across sessions, users, and models, and reports extraction through a weaker model rather than a direct jailbreak of the stronger one.
Two things were not taken. OpenAI is explicit that the operators did not break the encryption protecting the reasoning, did not access stored user conversations, and did not lift model weights. The target was the model's working memory, not the model itself or its users' data.
The disclosure matters beyond OpenAI. The company says it shared details with the Frontier Model Forum, an industry body whose members include the other major frontier developers, and the post itself states that the manipulation is not unique to OpenAI's models. Any lab that serves reasoning traces behind a public API faces the same replay pathway until the underlying mechanisms are changed.
OpenAI says it has closed the replay pathway for adversaries who already possess encrypted reasoning, rolled out stronger protections across users, workspaces, organizations, and model families, added streamed-output checks, and coordinated account-level enforcement with partners. The company is also explicit that further mitigation and propagation work across cloud partners continues, which means the attack surface is narrower than last month but not yet closed.
Two open questions sit underneath the disclosure. First, attribution. OpenAI says a core cluster was associated with individuals linked to Moonshot AI, a Chinese frontier developer, while cautioning that it is unclear whether all observed operators came from a single actor. That is an interested-party attribution, not an independent identification, and the post names no other party. Second, source independence. The campaign description, the headline numbers, and the mechanism narrative all come from OpenAI. The accompanying paper is co-authored by OpenAI researchers and was submitted on August 10, 2026; its full methodology and results have not been independently assessed here.
The disclosure puts a shape on a new attack class. Adversaries no longer need to breach a model lab to clone its reasoning. They can manipulate the model into producing the traces itself, at scale, through the same public surface everyone else uses. The next test is whether other frontier developers disclose comparable activity, and how quickly the mitigations described this week travel across the industry.