- On September 8, the NSA, FBI, and CISA issued a joint advisory, AA26-251a, accusing six China-based AI companies, DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI, of industrial-scale distillation of US frontier models.
- The advisory says the campaigns have run since at least late 2024 and target the Claude, GPT, Gemini, and Grok families, using native APIs, gray-market transfer stations, bulk premium subscriptions, and chain-of-thought extraction.
- US officials argue distillation lets Chinese models close the capability gap without paying the research bill, and assess that DeepSeek's often-cited $5.6 million training figure is misleading because it excludes what the labs spent copying US models.
NSA, FBI, and CISA name six Chinese labs in a joint distillation advisory
On September 8, three US agencies did something they had avoided for two years. In a joint cybersecurity advisory numbered AA26-251a, the NSA, FBI, and CISA named six specific China-based AI companies, DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI, and accused them of systematically extracting the capabilities of American frontier models to train their own. The agencies call it industrial-scale distillation, and they say it has been running since at least late 2024.
The mechanics described in the advisory are as much about procurement as about code. The labs are said to route requests through native APIs, remote cloud providers, and third-party aggregators, to lean on gray-market "transfer stations" that proxy around geographic restrictions, and to buy premium subscriptions in bulk and share them across teams. On top of that sits the technical layer, chain-of-thought reasoning extraction and automated failover that keeps the harvesting running when one pathway is cut. The four US model families named as targets are Claude, GPT, Gemini, and Grok.
The agencies were blunt about why it matters.
China-based companies' use of distillation against U.S. models enables Chinese AI models to continually close the technology gap without expending high research and development costs.National Security Agency, September 8, 2026
Why distillation attacks the one advantage US labs thought they had
The moat that frontier labs have counted on is cost. Training a leading model takes years of research, billions in compute, and talent that only a handful of companies can afford. Distillation attacks that moat directly, because it copies the expensive part, the model's behavior, without paying for the discovery. A student model learns from the teacher's outputs, and if the teacher is a paid API, the tuition is a subscription fee. This is why Washington's assessment that DeepSeek's $5.6 million training claim is misleading matters beyond one company. The point is not the exact figure, it is that the reported cost of Chinese frontier models has never included the American research it was built on top of.
This is not a new suspicion. The White House flagged distillation as a national-competitiveness issue in its moonshot review over the summer, and Santage has tracked the steady rise of Chinese models inside US enterprise token traffic. What changed on September 8 is specificity. Six companies are named, four target families are listed, and a two-year timeline is on the record from the agencies that watch foreign networks for a living.
The enforcement problem hiding inside the advisory
Naming the labs is the easy part. Stopping them is where the advisory quietly concedes the difficulty. Distillation through a public API looks, request by request, like ordinary paid usage. When those requests are laundered through aggregators, gray-market resellers, and bulk-bought premium seats, attribution collapses into probability rather than proof. Every US lab already bans distillation in its terms of service. The advisory is an admission that terms of service are not a defense when the buyer is willing to route around geography and rotate accounts faster than the provider can close them.
Every US lab already bans distillation in its terms of service. This advisory is an admission that a contract is not a firewall.
That leaves the frontier labs with two uncomfortable options. They can lock down what makes their models most useful, by hiding chain-of-thought traces and throttling the exact capabilities that enterprise customers pay for, which taxes legitimate users to slow illegitimate ones. Or they can accept that any capability exposed through an API is, over time, a capability that can be copied, and compete on the speed of the next model rather than the secrecy of the current one.
What this means for the export-control and open-weight debate
The advisory reframes the China AI question. For two years the policy conversation has been about chips, on the theory that denying Beijing the best hardware would slow its models. Distillation is a reminder that capability leaks through software channels that no chip ban touches. A lab that cannot buy an H200 can still buy a ChatGPT subscription, and the advisory says several have. It also complicates the open-weight argument. When labs like Abu Dhabi's MBZUAI release fully open frontier models, they hand over the weights deliberately, which at least makes the transfer honest and legible. The harder case is the closed model that is copied through its own front door, one metered request at a time.
Washington has now put names, methods, and a timeline on the practice the whole industry suspected. The advisory does not offer a fix, because there may not be a clean one. The uncomfortable conclusion is that the American lead in AI was always more perishable than a training budget suggested, and that the cost of defending it may be making the best models a little less open, a little slower, and a little less useful to everyone, in order to make them harder to steal.
Santage is committed to independent, transparent journalism. This article is produced in accordance with Santage's Editorial Standards and aims to provide accurate and timely information. Readers are encouraged to verify information independently.