News

OpenAI Says Moonshot-Linked Operators Used Encryption Trick to Steal Hidden AI Reasoning

OpenAI says operators tied to China's Moonshot AI sent about 16,000 prompts in two days to make its models decrypt their own hidden reasoning, in a distillation campaign it shut down by July 28.

OpenAI Says Moonshot-Linked Operators Used Encryption Trick to Steal Hidden AI Reasoning

OpenAI says a coordinated group tried to steal the hidden reasoning behind its models, and it pinned the core of the operation on people associated with Moonshot AI, the Chinese startup behind the Kimi chatbot. The attackers did not break any encryption. They talked the model into decrypting its own protected reasoning for them.

The disclosure puts a name on a fear the frontier labs have voiced for a year: that rivals are training cheaper models on the outputs of expensive ones. It also exposes a class of bug that OpenAI says may exist in other AI systems too.

A burst of 16,000 prompts

According to OpenAI, the activity started around July 1, 2026. It surged on July 24 and 25, when about 16,000 prompts from more than 4,000 users matched a single extraction pattern. Related activity spanned more than 15,000 users in total before OpenAI says it fully disrupted the campaign by July 28.

That scale is the signature of adversarial distillation. The goal is not one stolen answer but a large corpus of high-quality reasoning traces that can be used to teach a smaller model to think like a bigger one, at a fraction of the training cost.

The trick: decrypt it somewhere else

Reasoning models like OpenAI's produce a chain of internal steps before giving an answer. OpenAI keeps that raw reasoning hidden and can return it to clients in encrypted form, so it can be passed back into later requests without ever being readable.

The operators found a way around that. They copied the encrypted reasoning data from one conversation, opened a separate conversation and asked the model to decrypt and transcribe it in plain text. A bug allowed encrypted data created in one chat to be decrypted in another, so the model obliged.

OpenAI was careful to say what did not happen. The operators did not crack its encryption, did not compromise a database and did not gain direct access to stored user conversations. They manipulated normal model interactions so that protected reasoning was reproduced in a form the requester could see, which broke its terms of service rather than its cryptography. Outside researchers independently reported a similar vulnerability to OpenAI in August.

How strong is the Moonshot link?

OpenAI attributes a core cluster of the activity to individuals associated with Moonshot. It has not published hard technical evidence for that connection, such as infrastructure indicators or account details, so for now the attribution rests on OpenAI's word. Moonshot did not immediately respond to requests for comment.

The timing adds weight to the accusation without proving it. Only weeks earlier, Anthropic accused Chinese developers including Moonshot and Alibaba of secretly using Claude to help train rival systems. Two US labs pointing at the same company in quick succession turns distillation from a vague grievance into an industry-level dispute with a geopolitical edge.

The fix, and the warning

OpenAI says it fixed the bug that let encrypted data cross between chats, tightened signup and infrastructure controls, expanded network monitoring and banned the offending accounts. It also shared its findings through the Frontier Model Forum and with government channels.

The more uncomfortable part is OpenAI's own caveat that the same class of vulnerability may exist in other AI models. Any provider that hands clients an encrypted blob and later lets the model read it back has to make sure that blob cannot be replayed in the wrong context. It is the AI version of a classic confused deputy problem, where a trusted component is tricked into using its privileges on an attacker's behalf.

It is also another reminder that the model itself is now part of the attack surface. Cyberpresso has covered how OpenAI's newest systems are being pulled into supply chain attacks and how researchers have escaped its sandboxes through DNS. This time the asset being drained was not user data but the reasoning OpenAI spends billions to produce.

Cyberpresso: daily cyber & AI brief

Free daily newsletter, read in 5 minutes.

Subscribe free