Companies·3 min read
By BitsMindsSource: OpenAI

OpenAI Ties Reasoning-Theft Campaign to Moonshot AI

OpenAI says a coordinated campaign, its core traced to people associated with Kimi maker Moonshot AI, sent 16,000 extraction attempts from more than 4,000 accounts in two days to pull out the hidden reasoning of its models.

16,000 PROMPTS BITSMINDS.COM
Share:

OpenAI says it has broken up a coordinated attempt to copy the hidden reasoning of its models, and it is pointing at one of China’s best-known AI labs. In a report published on 30 September, the company attributes the “core cluster” of the activity to individuals associated with Moonshot AI, the Beijing developer of the Kimi models. The operators, OpenAI writes, “manipulated model interactions so that protected reasoning could be reproduced in forms visible to the requester,” in a coordinated way that broke its terms of service.

The timeline is short and sharp. The campaign started at low volume on 1 July, then spiked on 24 and 25 July to about 16,000 extraction attempts from more than 4,000 accounts. Following the prompt patterns outward, OpenAI found related activity across more than 15,000 users, and says the operation was fully shut down on 28 July. According to GovInfoSecurity, one technique was to copy the encrypted reasoning from one conversation and ask a model in a different conversation to decrypt and transcribe it. OpenAI stresses that nobody broke its encryption, breached a database or reached stored user chats; the attack worked through the products as designed, just at scale and in combination.

That approach lines up with academic work published in August. As The Hacker News reports, researchers from MATS, the ELLIS Institute Tübingen and Synk showed that a provider’s encrypted reasoning traces are interchangeable across sessions, users and even models, so a trace captured from a strong model can be fed to a weaker one that will then spell it out in plain text. Reasoning traces are prized because they show how a model works through a problem, which is exactly what a rival needs to train its own model to think the same way.

OpenAI says it has banned the accounts, tightened sign-up checks, added monitoring and hardened its hidden reasoning so it cannot be decrypted on request. It did not publish the technical evidence behind the Moonshot attribution, citing security. The company frames the issue as safety as well as business: extracted reasoning, it argues, could be used to train another model “without preserving the safeguards.” Moonshot did not answer requests for comment from The Register.

Moonshot has been here before. Anthropic has accused it of relaying customer requests to Claude and training on the answers, and a joint NSA, CISA and FBI advisory in September named it among six Chinese firms running mass distillation campaigns. The wider fight has escalated all year: Anthropic made similar claims about Alibaba’s Qwen, and Meta barred its engineers from using Claude Code and Codex for fear that rival outputs would leak into its training data. All of this comes as Moonshot’s Kimi K3 competes near the frontier as an open-weight model.

Not everyone is sympathetic. The Register ran the story under the headline “Irony alert,” noting that OpenAI built its models on scraped web data while now treating the copying of its own outputs as a security threat. The legal line between learning from a model’s answers and stealing them remains unsettled. What the disclosure does make clear is that hiding a model’s reasoning is now an arms race in its own right, and that a few thousand ordinary accounts were enough to put real pressure on it.

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.

Related Articles