Disrupting a coordinated model-distillation campaign
OpenAI attributed a core cluster of the extraction attempts involving over 15,000 accounts to individuals associated with Moonshot AI, the developer of Kimi.
- Operators manipulated model interactions and replayed encrypted reasoning across conversations to coerce models into decrypting hidden reasoning.
- The campaign peaked on July 24 and 25 with 16,000 attempted extraction requests from 4,000 users before full disruption on July 28.
- OpenAI deployed mitigations including banning accounts, closing reasoning replay pathways, and adding checks to hold streamed output exposing reasoning.
- The company shared threat findings with industry partners through the Frontier Model Forum and government channels.
AI developers face risks of model capability theft and safeguard bypasses through adversarial distillation, requiring defenses across reasoning replay artifacts and streaming outputs.

Sources
Read this as text
Back to the AI news