Open source is not open season: Washington threatens sanctions over Chinese AI firms' industrial-scale model distillation
A joint advisory from the NSA, FBI, and CISA accuses six Chinese AI labs — including DeepSeek, Alibaba, and Moonshot AI — of systematically siphoning capabilities from American frontier models. Treasury Secretary Scott Bessent says sanctions and Entity List designations are on the table — and the White House has now named Moonshot's Kimi K3.

For years, the AI contest between the United States and China has been fought over chips. This month the battleground moved to the models themselves.
In early September, three American national-security agencies — the NSA, the FBI, and the Cybersecurity and Infrastructure Security Agency — issued a rare joint advisory accusing six Chinese AI companies of running industrial-scale campaigns to strip capabilities out of American frontier models for their own systems. The technique at the center of the dispute is distillation, and the Treasury Secretary is now openly discussing sanctions. This week, the White House put a name and a number on the alleged operation: Moonshot AI, and Anthropic's Fable models.
What the agencies allege#
The advisory names six firms: DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI. Together, they allegedly pulled billions of tokens out of American models — Claude, GPT, Gemini, and Grok variants — in operations dating back to at least late 2024, "likely" with Beijing's awareness. The language is unusually blunt: the advisory calls distillation not a side trick but "the critical core" of the companies' development programs.
The agencies describe a playbook that broke the rules at every step: fraudulent accounts, proxy networks the advisory calls "transfer stations," geographic restrictions routed around, and the American labs' terms of service ignored. Then it gets specific. DeepSeek allegedly targeted reasoning capabilities and specialized optimizations from GPT-4, GPT-5, and several Claude versions to build its R1 and V3 models — and the agencies pointedly note that DeepSeek's famous $5.6 million training-cost figure leaves out the cost of the data it is said to have lifted. Moonshot AI is accused of feeding on Claude to build Kimi K3, while Alibaba, MiniMax, StepFun, and Z.AI allegedly harvested capabilities for software engineering, coding, and customer-service functions.
Why distillation cuts so deep#

Distillation is an ordinary engineering technique: a large "teacher" model answers millions of prompts, and a smaller "student" trains on those answers, inheriting much of the teacher's ability at a fraction of the cost. The line Washington is drawing is about consent and scale. Bombard someone else's model with automated queries to copy its behavior and you are no longer doing research — you are strip-mining their investment.
The economics explain the heat. Training a frontier model can cost hundreds of millions of dollars; distillation can compress years of research into weeks of API bills. If a rival can harvest a flagship model's reasoning for pennies per token, the moat that justifies the spending collapses — and with it, the assumption that American compute investment buys a durable advantage.
The escalation this week#

Treasury Secretary Scott Bessent escalated first. In an early-September post on X, he warned that "covert, industrial-scale distillation attacks" crossing into intellectual-property theft would put "sanctions and Entity List designations on the table." On Wednesday he sharpened it: Washington supports open-source AI — but "open source is not open season on American IP."
The specifics arrived from the White House. Science chief Michael Kratsios posted that US intelligence indicates Moonshot AI used Anthropic's Fable models to develop Kimi K3 through large-scale distillation — and that Moonshot had reportedly acquired or accessed servers equipped with NVIDIA GB300 chips, hardware squarely inside Washington's export-control crosshairs. It echoes a July warning from the White House science office, which accused Moonshot of building a distillation platform and obtaining banned Nvidia chips. Anthropic, for its part, has alleged that three Chinese labs generated more than 16 million exchanges with Claude through roughly 24,000 fraudulent accounts.
Beijing pushes back — and Washington argues with itself#
Beijing rejects all of it. Foreign Ministry spokesperson Mao Ning said China's AI advances reflect its own self-reliance and strength in science and technology, and told Washington to stop making "unfounded accusations" and smearing China. The Chinese Embassy in Washington called the American framing "a deliberate attack on China's development and progress in the AI industry."
But the sharpest friction is inside Washington itself. Representative Ted Lieu, a California Democrat, publicly questioned the administration's coherence: threatening sanctions over stolen AI know-how while still permitting sales of high-performance AI chips to China looks like arming both sides of the contest. NVIDIA chief executive Jensen Huang took the opposite tack, saying American companies should "absolutely" use Chinese AI models and calling distillation "fundamental to intelligence" — across much of the industry, the technique remains as ordinary as compression.
What to watch#
Three things now matter more than the rhetoric. First, timing: the accusations land weeks before President Trump is expected to host China's Xi Jinping in Washington, with AI on the agenda — the advisory either sets the table for negotiation or poisons it. Second, enforcement: Bessent's "on the table" only bites if names actually land on the Entity List, which would choke a listed firm's access to American technology overnight. Third, the industry's own defenses: the agencies urged US labs to watch for distillation behavior, quietly alter model outputs when they suspect it, and share threat intelligence with competitors — an ask that runs against every competitive instinct in Silicon Valley.
The era in which model weights were the asset is giving way to something stranger: one in which the asset is the model's behavior — and behavior can be copied one API call at a time.