Key Points
- OpenAI says it disrupted a coordinated campaign to extract its models' protected reasoning in July.
- The company attributes a core cluster of the activity to individuals associated with Moonshot AI, Kimi's developer.
- OpenAI frames adversarial distillation as a safety and national security risk shared across the industry.
The latest:
A coordinated effort to pull protected reasoning out of OpenAI’s models was identified and shut down, the company said, with the earliest activity dated to the first week of July. OpenAI said it attributes a core cluster of that activity to individuals associated with Moonshot AI, the developer of Kimi. It added that no encryption was broken and no database was compromised.
Details:
- The technique: OpenAI said operators manipulated model interactions so protected reasoning could be reproduced in forms visible to the requester, at scale and against its terms of service. Protected reasoning is the model’s internal record for working through a task, which can expose information withheld from the final answer.
- One method observed: According to OpenAI, operators copied encrypted reasoning from one conversation, then asked a model in a separate conversation to decrypt and transcribe the hidden content. Independent security researchers separately flagged related cross-model and conversation-compaction weaknesses through responsible disclosure.
- The numbers: OpenAI said activity began July 1 at low volume before high-volume spikes on July 24 and 25, totalling 16,000 requests using a relevant extraction pattern from more than 4,000 users. Further investigation identified related prompt-pattern activity across a cluster of over 15,000 users.
- The timeline: The company said it fully disrupted the identified cluster by July 28, roughly four weeks after the first observed activity. OpenAI did not say when the campaign was first detected internally, nor what triggered the investigation that mapped the wider user cluster.
- The attribution caveat: OpenAI said it is unclear whether all operators observed during the period originated from a single actor, and limited its attribution to a core cluster tied to individuals associated with Moonshot AI. It did not say whether it had contacted Moonshot AI or whether legal action is planned.
- What was ruled out: OpenAI stated the operators did not break its encryption, compromise a database, or gain direct access to stored user conversations. It also said the manipulation is not a vulnerability unique to its own models, and that similar techniques may affect other advanced AI systems.
- The response: OpenAI said it banned or restricted fraudulent accounts, tightened signup and infrastructure controls, and expanded monitoring for related networks. It also strengthened protections for hidden reasoning across users, workspaces, organizations and model families, and shut a route that let anyone holding another user’s encrypted reasoning replay it and recover the contents.
- The stated risk: OpenAI argued extracted reasoning could be used to train another model without preserving the safeguards applied to the original model’s user-facing outputs. At scale, it said, distillation can accelerate the transfer of advanced capabilities without matching investment in safety. Findings were shared through the Frontier Model Forum and government information-sharing channels.
Background:
Adversarial distillation refers to using one model’s outputs or reasoning, without authorization, to help train, reproduce or improve another model. The Frontier Model Forum, cited by OpenAI as its industry channel here, is an industry body for coordination on frontier AI safety and security.
Between the lines:
OpenAI’s own framing does the heavy lifting: it names a competitor’s associates while conceding it cannot say all the activity came from one actor. That gap between a named core cluster and 15,000 flagged users is where the attribution’s limits sit. By stressing that the technique is not specific to its models and routing findings to industry and government channels, the company casts a commercial dispute as a sector-wide security problem.
What’s next
Watch for any response from Moonshot AI, whether OpenAI pursues enforcement beyond account bans, and whether other Frontier Model Forum members disclose similar extraction attempts against their own models.