OpenAI has disclosed a coordinated campaign that attempted to extract hidden reasoning from its AI models, marking a significant security challenge for the company. The activity, which began on July 1, peaked on July 24 and 25, with 16,000 matching requests originating from more than 4,000 users, according to the company's incident report. By July 28, OpenAI had disrupted the broader cluster of attacks, though the public disclosure came on September 30.
Attack Methodology
The attackers did not breach encryption or directly access model databases. Instead, they manipulated model conversations to coax protected reasoning into readable output. One notable technique involved copying an encrypted reasoning block from one conversation and inserting it into another, then asking the second model to decode it. This method exploits the handoff process where reasoning models generate internal work before producing a final answer, and providers often conceal that work to reduce data leakage and hinder model copying.
Encrypted Reasoning Vulnerability
Some APIs return hidden reasoning as encrypted text, which is sent back to the model when a conversation continues. While this prevents clients from reading the block directly, providers' models still need to recover it. The portability of these blocks across sessions, users, or model variants creates a vulnerability. An attacker can search for a weaker decoding path within the same ecosystem.
An independent security paper documented this broader attack class across OpenAI, Anthropic, and Google systems. Researchers moved encrypted traces into weaker models from the same provider and induced plaintext output. They also decoded 315,320 blocks found in public repositories, exposing 367 personal-data artifacts and 182 credentials, as reported in the paper.
OpenAI's Response
OpenAI has confirmed the researchers' attack paths and has tightened the binding of hidden reasoning across users, workspaces, organizations, and model families. The company also closed a replay route involving another user's encrypted reasoning. New output checks are now in place to hold streamed responses that appear to reveal hidden work.
Attribution and Scale
OpenAI attributes a core cluster of the attacks to people associated with Moonshot AI, but notes that the observed operators may not belong to a single actor. The public post lacks technical attribution evidence and does not name individual accounts or the targeted models. The scale figures also require careful interpretation: they describe attempted extractions, not successful ones. OpenAI has not disclosed a success rate, and the average of fewer than four requests per user at peak highlights the difficulty of detecting coordinated networks through per-account rate limits alone.
Anthropic's September threat report separately alleges that Moonshot relayed nearly 300,000 requests over ten days through 5,380 accounts, using a cross-session replay technique with Claude reasoning signatures. While this technical overlap supports the plausibility of the attack, it does not independently prove OpenAI's attribution.
Implications for AI Security
Distillation, a standard training method where a larger teacher model generates examples for a smaller student, is not inherently problematic. The security issue here is covert extraction without permission, often using false accounts or exploiting replay flaws. The disclosed evidence concerns platform abuse rather than a formal legal finding.
Encryption alone is insufficient; the encrypted reasoning blocks are treated as recoverable model context, making them portable bearer objects. Stronger designs bind each block to a single tenant, session, model, and purpose, with short expiry, replay detection, and server-side state to narrow reuse. Cryptographic binding would render copied blocks unusable outside their original context, turning many replay attempts into cheap validation failures. Providers also need network-level detection across accounts to counter scaled campaigns.
Partner-hosted models introduce additional boundaries, as first-party API protections may not cover every cloud route. Tool outputs create another channel beyond visible text. OpenAI acknowledges that both areas require further work.
Comparing an API credential—which should have narrow permissions, a short life, and a clear owner—encrypted reasoning needs similar security scope. The decoding model is effectively a privileged service; if it accepts a block from the wrong context, encryption only preserves secrecy until the next model call.
Next Steps
The next verifiable milestone is an independent retest across first-party and partner-hosted deployments. Researchers should examine whether the fixes hold under varied conditions. As AI models become more integrated into financial and enterprise systems, the security of their internal reasoning processes will be critical to maintaining trust and data integrity.



