Industry·5 min read
By BitsMindsSource: CISA / NSA / FBI

Feds Accuse Six Chinese AI Firms of Mass Distillation

A joint NSA, CISA and FBI advisory published September 8 names DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z.AI as running "industrial-scale" distillation campaigns against Claude, GPT, Gemini and Grok since late 2024 — billions of tokens pulled through fraudulent accounts, grey-market proxies and jailbreak prompting. The agencies also argue DeepSeek’s famous $5.6M training cost leaves out the data it took.

JOINT CYBERSECURITY ADVISORY NSA · CISA · FBI AA26-251A SEPT 8, 2026 Claude GPT Gemini Grok DeepSeek · Moonshot AI · Alibaba · MiniMax · StepFun · Z.AI Billions of tokens · millions of requests · since late 2024 BITSMINDS.COM
Share:

The National Security Agency, the Cybersecurity and Infrastructure Security Agency and the FBI published a joint cybersecurity advisory on September 8 accusing six China-based AI companies of running "industrial-scale knowledge distillation campaigns" against American frontier models. The advisory, catalogued as AA26-251A, names DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z.AI, and says the extraction has been running since at least late 2024.

Distillation is an ordinary technique in machine learning: query a strong model, collect its outputs, and train a smaller model to imitate them. What the agencies describe is that technique industrialised and pointed at competitors' paid APIs. The campaigns pulled "billions of tokens across millions of exchanges", the advisory says, from Claude variants including Sonnet, Opus, Haiku and Fable, from OpenAI's GPT-4, GPT-4o and GPT-5 series, and from Gemini and Grok models.

The per-company detail is where the advisory is most specific. DeepSeek is alleged to have distilled Claude, Gemini, GPT and Grok to generate synthetic training data for R1 and later R3. Moonshot AI is said to have extracted from 18 different US models while building Kimi K2 and K3, with millions of exchanges logged from mid-2025. Alibaba is placed against Claude and GPT-5 in late 2025 to improve the Qwen family; MiniMax against Claude Code specifically, harvesting chain-of-thought traces and reinforcement-learning data; StepFun across reasoning and coding models from late 2025 into early 2026; and Z.AI against GPT-5.5 and Claude Opus, reaching billions of tokens by mid-2026.

The tradecraft described goes well beyond signing up for an API key. The agencies document tens of thousands of fraudulent accounts run simultaneously, a grey market of proxy services the advisory calls "transfer stations" used to defeat regional restrictions, laundering of traffic through third-party API aggregators and cloud providers, automated failover between access routes when one is blocked, metadata sanitisation to strip organisational identifiers, and jailbreak prompting aimed at making models reveal reasoning they would normally keep hidden. MiniMax is additionally accused of attempting prompt-injection attacks to pull chain-of-thought out of Claude Code.

One line in the advisory is aimed squarely at the most repeated number in the industry. DeepSeek's widely cited $5.6 million training cost, the agencies argue, "excludes the cost of data acquired through malicious distillation" — a claim that, if accepted, unwinds the efficiency story that moved markets when R1 landed in early 2025. More broadly, the agencies write that these campaigns "form the core — not merely a supplement" of how the named companies build models.

None of this is entirely new ground for the sector. Anthropic publicly accused Alibaba of mass Claude distillation in June, and BitsMinds covered evidence in July that Chinese military labs had distilled GPT-3.5 and Claude 3 Haiku with no export control covering the practice. What changed on Monday is the source: this is the first time the US intelligence and law-enforcement establishment has put its own name to the attribution, named six commercial firms in one document, and mapped which American model each allegedly targeted.

The mitigations section is unusually candid about how thin the defences are. The agencies advise model providers to watch subscription-to-usage ratios, accounts that hit maximum throughput immediately after registration, and enterprise-scale traffic on consumer plans — and, more aggressively, to alter responses to high-confidence distillation requests rather than simply blocking them, by "reducing reasoning depth" or "presenting correct information with different reasoning". That last suggestion amounts to poisoning the well for a suspected scraper, and it only works if providers, cloud platforms and API aggregators share detection signals with one another, which the advisory also recommends.

It is worth being precise about what the document is and is not. An advisory is guidance, not an indictment: AA26-251A carries no sanctions, no export-control designation and no enforcement action, and the six named companies had not publicly responded at the time of publication. Nor do the agencies publish the underlying telemetry, so the specifics rest on attribution the reader is asked to take on trust. What it does establish is a government position — and a paper trail that any future licensing rule, procurement ban or Section 337 complaint can be built on top of.

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.

Related Articles