METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

OpenAI says it disrupted Moonshot-linked distillation campaign

OpenAI said it disrupted a coordinated distillation campaign aimed at extracting protected reasoning from its models, and attributed a core cluster to individuals associated with Moonshot AI, the developer of Kimi. Over two days in July, more than 4,000 users sent 16,000 requests.

OpenAI says it disrupted Moonshot-linked distillation campaign

Image: METAL

Summary

  • OpenAI announced on September 30 that it identified and disrupted a coordinated distillation campaign designed to extract protected reasoning from its models, attributing a core cluster to individuals associated with Moonshot AI.
  • The activity began on July 1 and spiked on July 24 and 25 with 16,000 requests from more than 4,000 users; related patterns spread across more than 15,000 users but were fully disrupted by July 28.
  • OpenAI responded with account bans, stronger reasoning protections, and information sharing through the Frontier Model Forum and government channels, and framed the matter as a terms of service violation rather than an issue with open models.

OpenAI announced on September 30 that it had identified and disrupted a coordinated campaign to extract protected reasoning from its models. The company attributed a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi. According to reports, this is the first time OpenAI has publicly named Moonshot. The operators did not break encryption or break into a database; instead they manipulated interactions with the model so that hidden reasoning was reproduced in a form visible to the requester.

OpenAI characterized the activity as adversarial distillation. According to the announcement, adversarial distillation is the systematic and unauthorized use of one model's outputs or reasoning to help train, reproduce, or improve another model. Protected reasoning is the internal record a model keeps while working through a task; it can reveal information withheld from the final answer and can help others reproduce the model's capabilities. The company said in the announcement that "this manipulation is not a vulnerability unique to OpenAI's models."

The technique was new. The operators copied encrypted reasoning from one conversation and then asked a model in another conversation to decrypt and transcribe that hidden reasoning. Independent security researchers also flagged cross-model and conversation-compaction vulnerabilities through responsible disclosure, and OpenAI confirmed that the attack paths they identified were real.

The scale emerged date by date. The activity began at low volume on July 1, and on July 24 and 25 more than 4,000 users sent 16,000 requests using an extraction pattern in a burst. Further investigation found the related prompt pattern spread across a cluster of more than 15,000 users, which OpenAI fully disrupted by July 28. In a footnote, the company said these figures describe attempted, not necessarily successful, extractions. It also wrote that it is unclear whether all the operators observed came from a single actor.

The response had three strands. OpenAI banned or restricted fraudulent accounts, strengthened signup and infrastructure controls, and expanded monitoring for related networks. It closed a pathway that allowed someone who already held another user's encrypted reasoning to replay it and recover its contents, and added checks to detect and hold streamed output that might expose reasoning. When the activity moved through third-party services, it worked with those providers to find and disrupt the accounts involved. It passed its findings to the Frontier Model Forum and government information-sharing channels.

OpenAI drew the line at its terms of service. According to reports, Caroline Zier, who leads strategic national security policy initiatives at OpenAI, said, "Our concern is about violation of our terms of service, not open models or legitimate distillation." At the same time, the announcement flagged safety and national security risks, noting that extracted reasoning could be used to train another model without the original model's safeguards and that large-scale distillation could become a channel for transferring advanced capabilities without the same investment in safety.

From a legal perspective, the weight of this announcement rests on contract and evidence. What OpenAI took issue with was not hacking but large-scale use in a way its terms prohibit, and the evidence it offered was an operational record of dates, request counts, and user counts. The attribution was also worded carefully. The company assigned only the core cluster, not the whole campaign, to individuals linked to Moonshot, and chose to name individuals associated with the company rather than Moonshot itself. A terms of service violation can be answered with account bans and access blocks, but the tools for holding a developer outside the country accountable still rely on industry information sharing and government channels.

This is not the first time Moonshot has been named. METAL previously reported on Anthropic's report disclosing unauthorized distillation by seven Chinese AI labs, and according to reports Anthropic had earlier made a similar claim against Moonshot. According to reports, Michael Kratsios, director of the White House Office of Science and Technology Policy, said in late July that the US government had "information" that Moonshot distilled Kimi K3 from Anthropic's models, and a US advisory this month named six Chinese companies and the US models each one targeted. China rejected the advisory as unfounded accusations.

There is pushback as well. According to reports, David Sacks, who oversaw AI policy in the Trump administration, has criticized such reports as a push to get the US to ban rival open models. Moonshot is also under pressure at home, with reports this month that Chinese regulators are investigating DeepSeek and Moonshot.

The OpenAI announcement that METAL reviewed describes this action as a beginning rather than an end. The company expects adversarial distillation attempts to become more sophisticated as frontier models improve, and said partner-hosted deployments need the same protections as its own services and that attacks targeting tool output require defenses that examine more than ordinary visible text. It also warned that systems supporting portable or replayable reasoning artifacts may face similar risks. Its focus from here is threefold: stronger technical protections against extraction, better detection and enforcement against coordinated campaigns, and deeper threat-information sharing between industry and government. The task this case leaves the industry is that hiding reasoning is not enough; systems must be designed with an eye to where hidden reasoning can flow.

Comments