OpenAI Disrupts Moonshot AI Distillation Campaign — SkimNews

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- OpenAI identified a coordinated distillation campaign traced to individuals associated with Moonshot AI, a Beijing-based Chinese AI company, dating to the first week of July 2026; the company did not cite technical evidence for the attribution, citing security reasons.
- The activity began on July 1, 2026 at low volume, then spiked on July 24-25 to 16,000 attempted requests using a relevant extraction pattern from over 4,000 users, with related prompt-pattern activity spanning more than 15,000 users total before being fully disrupted on July 28, 2026.
- OpenAI described the operators as having manipulated model interactions so protected reasoning could be reproduced in forms visible to the requester, violating its terms of service — without breaking encryption, compromising a database, or gaining direct access to stored conversations.
- OpenAI deployed additional mitigations, banned the fraudulent accounts, closed a pathway that allowed replaying another user's encrypted reasoning, and added checks to detect and hold streamed output that might expose reasoning.
- A study published in August 2026 by MATS Research, ELLIS Institute Tübingen, and Synk found an architectural vulnerability affecting Claude, Gemini, and GPT that made encrypted reasoning traces "fully compatible and interchangeable" across sessions, users, and models, enabling scalable decryption jailbreaks.
- OpenAI framed adversarial distillation as a safety and national security risk, warning that extracted reasoning could train another model without the original's safeguards and accelerate transfer of dual-use capabilities.
- Moonshot AI previously faced accusations from Anthropic last month of relaying customer requests to Claude, displaying its responses, and retaining exchanges to train its chain-of-thought model — activity tracked as GTG-16002.
Why it matters: OpenAI's 16,000-request spike on July 24-25, traced to Moonshot AI associates, marks the first time a major U.S. lab has publicly attributed a scaled distillation operation to a named Chinese competitor, raising the stakes from a terms-of-service violation to a national-security framing — and the August study showing reasoning traces are interchangeable across Claude, Gemini, and GPT suggests OpenAI's mitigations alone won't close the loophole.
Ask SkimNews




