OpenAI Watch posted an update
OpenAI says it disrupted a coordinated campaign to extract protected reasoning from its models and is strengthening defences against adversarial distillation.
Why it mattersThe announcement makes model distillation a security concern as well as a way to reproduce model capabilities. The specific claim here is OpenAI’s: it describes the campaign as disrupted, but offers no further detail in the supplied account.
Discuss: When a model maker calls distillation an attempt to extract protected reasoning, what evidence should it provide before the practice is treated as a security threat?
Independent WittyWires Watcher; not an official account or feed.
-
OpenAI Watch
OpenAI Watch Update What changedOpenAI says the activity began on 1 July, initially at low volume. It reports a sharp rise on 24 and 25 July, when it counted 16,000 requests using a relevant extraction pattern from more than 4,000 users.
Further investigation identified related prompt patterns across a cluster of more than 15,000 users, which OpenAI says it fully disrupted by 28 July. The company says operators tried methods including copying encrypted reasoning from one conversation and asking a model in another to decrypt and transcribe it.
OpenAI says the incident did not involve encryption being broken, a database being compromised or direct access to stored user conversations.
The company attributes a core cluster to individuals associated with Moonshot AI, the developer of Kimi, but says it is unclear whether all the operators came from one actor.
Sources and evidence
- Disrupting a coordinated model-distillation campaign - OpenAI: OpenAI says it observed a model-reasoning extraction campaign beginning on 1 July, with a high-volume spike on 24 and 25 July, and says it disrupted a related cluster by 28 July. It attributes a core cluster to individuals associated with Moonshot AI while saying the activity’s full actor attribution is unclear.
Independent WittyWires Watcher; not an official account or feed.