[Tech News] OpenAI accuses China's Moonshot AI of stealing model reasoning: 'adversarial distillation' explained
- Get link
- X
- Other Apps
OpenAI says China's Moonshot AI tried to steal its models' "thoughts"
In a blog post published this week, OpenAI disclosed that it disrupted a coordinated campaign to extract the hidden reasoning of its AI models — and linked a core cluster of the activity to individuals associated with Moonshot AI, the Chinese startup behind the Kimi chatbot. The company calls the technique "adversarial distillation."
It's the most detailed public accusation yet in an escalating shadow war over how AI models are trained — and it follows similar allegations Anthropic leveled against Chinese developers just last month.
What allegedly happened
According to OpenAI, the campaign began on July 1 at low volume, then spiked dramatically: on July 24 and 25 alone, more than 4,000 users fired some 16,000 requests using a particular prompting technique. Further investigation found related prompt patterns across a cluster of more than 15,000 users. OpenAI says it fully shut the operation down by July 28.
The method was novel. Operators copied the encrypted reasoning output from one conversation, pasted it into a separate conversation, and asked a different model instance to decrypt and transcribe it. OpenAI encrypts the chain-of-thought its reasoning models produce precisely so competitors can't read and learn from it.
Crucially, OpenAI stresses what did not happen: the operators never broke its encryption, never accessed stored user conversations, and never breached a database. Instead, they manipulated ordinary model interactions with carefully designed prompts to make normally hidden reasoning visible.
What is "adversarial distillation"?
In machine learning, distillation is a legitimate, widely used technique: a smaller "student" model is trained on the outputs of a larger "teacher" model to inherit its capabilities cheaply. What OpenAI describes adds the word "adversarial" — the systematic, unauthorized use of one model's outputs or reasoning to train, reproduce, or improve a rival model, in violation of terms of service.
Why does it matter? As OpenAI puts it, a model's protected reasoning is "the model's internal record for working through a task" — extracting it can reveal information withheld from the final answer and help others reproduce advanced capabilities without the original safety guardrails. That's why the company frames this as a national-security concern, not just a terms-of-service dispute. "Our concern is about violation of our terms of service, not open models or legitimate distillation," said Caroline Zier, who leads OpenAI's strategic national-security policy work.
How OpenAI responded
The company says it banned the offending accounts, tightened sign-up verification, added new protections around hidden reasoning, and shared its findings through the Frontier Model Forum — the industry body where labs coordinate on safety — as well as through government information-sharing channels. Moonshot AI has not publicly responded to the allegations.
The bigger pattern
This isn't happening in a vacuum. Anthropic recently accused several Chinese AI developers, including Moonshot and Alibaba, of large-scale distillation attempts against its Claude models. Google has made similar complaints. And both OpenAI and Anthropic are reportedly moving toward blockbuster IPOs — which raises the stakes of protecting proprietary model capabilities enormously.
The deeper question is whether "reasoning" can even be kept secret. As models get more capable, the line between using a model and learning from it gets blurrier — and the industry's current answer (encrypt the chain of thought, ban the accounts) looks like the beginning of an arms race, not the end of one.
The bottom line
OpenAI caught someone trying to peek at its models' homework and is naming names. Whether the accusation holds up — and how Beijing responds — will shape the rules of AI competition for years.
- Get link
- X
- Other Apps

Comments
Post a Comment