Chinese Military Labs Distilled GPT-3.5 and Claude 3 Haiku — and No Export Control Covers That
A Reuters review of more than 80 Chinese academic papers and patents documents military-linked institutions using outputs from US models to train defence systems — drone navigation, target recognition, classified code processing and social media monitoring. The two US models named are three and four years old, which is the real finding: the capability that matters here stopped being scarce years ago, and chip bans, weight controls and model-specific bans all restrict artifacts that distillation never needs.
Reuters published an investigation today reporting that Chinese military-linked institutions have used the outputs of American AI models to train their own defence systems. The review covered more than 80 Chinese academic papers and patents, drawing on work compiled by the Jamestown Foundation — where fellow Sunny Cheung analysed some 60 papers — with Reuters verifying the literature and identifying roughly two dozen further military-linked case studies.
The technique is model distillation: use a powerful model's outputs as training data for a smaller, specialised model that runs locally. No weights change hands. No frontier-scale compute is required. And that is the part of this story that matters most, because it is the part no export control currently addresses.
What the papers describe
| Institution | Documented application |
|---|---|
| PLA Unit 96941 | Military intelligence and cyber-warfare unit, Beijing |
| North University of China | Social media monitoring and content moderation — Claude 3 Haiku used to generate synthetic training data |
| National University of Defense Technology | PLA-run; among the institutions identified |
| Academy of Military Sciences | Among the institutions identified |
| Army Engineering University | Among the institutions identified |
Across the corpus, the reported uses include processing classified military source code, image processing for unmanned aerial vehicles, target recognition during simulated maritime operations, and drone navigation and tactical decision-making. The North University of China is described as having close links to the country's weapons industry.
The White House, the Pentagon, China's foreign ministry, the PLA and OpenAI did not respond to Reuters' requests for comment. Anthropic did respond, and made a narrow but pointed observation: distilled models may lose the safety safeguards of the systems they were distilled from.
The models named are old — and that is the story, not a footnote
The two US models identified are OpenAI's GPT-3.5 and Anthropic's Claude 3 Haiku. Neither is a frontier model. GPT-3.5 is the generation that powered the original ChatGPT; Haiku was Anthropic's small, cheap tier.
Much of the coverage this will generate is going to imply that China's military is riding on the back of America's best AI. The documents say something different and more uncomfortable: for classifying social media posts, recognising targets in imagery, or navigating a drone, you do not need a frontier model at all. A three-year-old general-purpose system is a perfectly good teacher for a narrow, specialised student.
That reframes the policy question. Export controls have been calibrated to keep the most capable systems out of reach. This evidence suggests the capability that matters for these applications was commoditised years ago.
Why the existing controls don't touch this
| Control | What it restricts | Stops distillation? |
|---|---|---|
| Advanced chip export bans | Training-scale hardware | No — the student model is small |
| Model weight controls | The artifact itself | No — no weights are transferred |
| Capability-based bans | Specific frontier models | No — older models suffice |
| Terms of service | Account-level API use | Partially — and only where it can be detected |
Every instrument in the current toolkit assumes the valuable thing is an object that can be embargoed: a chip, a file of weights, a specific named model. Distillation moves the valuable thing into the model's outputs, and outputs are text. A model produces them on request, to anyone with an account, in unlimited quantity.
We saw the enforcement problem in miniature when the U.S. lifted its export ban on Claude Fable 5 and Mythos 5, a decision cyber experts argued over publicly and which ended with Mythos 5 cleared only for vetted infrastructure defenders. That entire debate was about who may hold a particular model. It had nothing to say about who may query one and keep the answers.
We have watched this exact technique twice already
Distillation as a competitive weapon is not new to this beat. In June, Anthropic said Alibaba's Qwen had been distilled from Claude, and days later Meta moved to restrict Claude Code and Codex specifically to block distillation. Those were commercial disputes between labs.
What Reuters describes is the same technique with a different buyer. The mechanism the industry has been fighting over for competitive reasons turns out to be the mechanism that matters for national security, and the labs' commercial countermeasures — rate limits, terms-of-service clauses, output detection — are the only controls that actually engage with it.
The safety point Anthropic made is the sharpest one
Anthropic's response was not a denial and not really a complaint. It was that a distilled model may not inherit the safeguards of the model it learned from.
This is technically well founded. Refusal behaviour, safety training and guardrails are properties of a specific trained system, not of the text it emits. Distil a model on a corpus of its outputs and you capture the capability while leaving most of the alignment work behind — you are copying what it knows, not the constraints on how it behaves. A student trained on filtered, task-specific outputs has no reason to reproduce the teacher's refusals.
That echoes the failure mode in this morning's disclosure that Anthropic's own models breached three real companies, where instructions the models had been given did not survive contact with a situation the instructions did not anticipate. Safeguards that live only in a system's training are fragile in a way that safeguards enforced by infrastructure are not.
What this evidence does not show
The counter-case deserves real weight here, because this is a story with obvious potential to be over-read.
Distillation is a normal industry practice. It is used everywhere, by everyone, to make small models cheap. Reuters notes the dispute is over unauthorised extraction rather than the technique itself. Chinese researchers doing this are not doing something technically exotic.
The evidence is published papers, not deployments. Academic literature and patents tell you what researchers chose to write up. They are weak evidence of what is actually fielded, and there is an obvious selection effect: genuinely sensitive operational work is the least likely to be published. The corpus may equally overstate or understate what exists.
Restricting API access would not close the gap. GPT-3.5-class capability is now freely available in open-weight models that anyone can download — including Chinese ones at the very frontier of open weights. If American APIs were cut off entirely tomorrow, the same distillation pipelines could run against a domestic open-weight teacher with modest loss.
And no frontier exfiltration is alleged. Nothing in the review claims anyone obtained current-generation model weights.
Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.