Days after claiming to have disrupted a Chinese botnet network that was allegedly targeting American critical infrastructure, the US has accused Chinese AI companies of ‘copying’ proprietary functionalities and capabilities of US AI companies. America’s Cyber Defense Agency, Cybersecurity & Infrastructure Security Agency (CISA), said in a release issued on 8th September that Chinese AI companies are conducting “industrial-scale knowledge distillation campaigns” to extract proprietary capabilities from leading AI models.
“China-based artificial intelligence (AI) companies are conducting systematic extraction of proprietary functionalities and capabilities of U.S. AI companies’ models through industrial-scale knowledge distillation campaigns that form the core—not merely a supplement—of their AI development strategy,” the press release stated.
The US authorities acknowledged that distillation is a legitimate and useful technique in AI research. However, the US officials claimed that China-based AI companies are engaging in “aggressive, malicious, and targeted distillation activities at an industrial scale that extract restricted proprietary functionalities and capabilities of U.S. frontier AI models.”
DeepSeek, Moonshot AI, Alibaba and other Chinese AI companies with CCP awareness are extracting capabilities of top US AI models like Claude, GPT, Gemini, and Grok
US officials have alleged that top Chinese AI models DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI extracted “billions of tokens across millions of exchanges/requests” from US frontier AI models, including variants of Claude, GPT, Gemini, and Grok, since at least late 2024. These alleged distillation campaigns are allegedly being conducted with Chinese government awareness.
The US authorities alleged that since 2024, DeepSeek has conducted organised distillation campaigns targeting reasoning capabilities, specialised optimisations, and domain-specific functions to train its R1 and V3 models.
Similarly, Alibaba allegedly leveraged industrial-scale distillation to improve its Qwen family of AI models.
The US authorities claimed that top Chinese AI models, including Moonshot AI, MiniMax, Stepfun, and Z.AI, also engaged in “malicious knowledge distillation of U.S. AI companies’ models.”
“Likely with Chinese government awareness, DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI extracted billions of tokens across millions of exchanges/requests from U.S. frontier AI models, including variants of Claude, GPT, Gemini, and Grok, since at least late 2024. DeepSeek has conducted organized campaigns since at least 2024 targeting reasoning capabilities, specialized optimizations, and domain-specific functions to train its R1 and V3 models. Alibaba leveraged industrial-scale distillation to improve the company’s Qwen family of AI models. Moonshot AI, MiniMax, Stepfun, and Z.AI also engaged in malicious knowledge distillation of U.S. AI companies’ models,” the US officials said.

The modus operandi
Explaining the modus operandi of the Chinese distillation campaigns targeting US frontier AI models, the US authorities said that several top Chinese AI companies route distillation requests via multiple pathways to obtain unauthorised access and violate the terms of the targeted US AI companies.
These pathways include native application programming interfaces (APIs), remote cloud providers, and third-party aggregators that automatically obfuscate user metadata to avoid detection.

In addition, the China-based AI companies allegedly use a gray market of proxies known as “transfer stations” to bypass geographic restrictions placed by US AI companies.

Through bulk procurement of the U.S. AI companies’ premium subscriptions shared across teams of developers, several China-based AI companies secure cost savings for their industrial-scale distillation campaigns.
Various Chinese AI companies have allegedly conducted prompt injection techniques against large language models (LLMs) by inserting prompts specifically designed for jailbreaking.
The advanced distillation methods include chain-of-thought (CoT) reasoning extraction, automated failover between pathways during blocking attempts, and sophisticated quality evaluation frameworks to detect defensive countermeasures. This tactic was used by DeepSeek.
Notably, the CoT data teaches student models not only factual knowledge but reasoning methodologies for complex agentic tasks, coding challenges, and logical proofs.
According to the US authorities, these Chinese AI entities use aggressive and adaptive discovery to systematically identify valuable extractable data.
With the help of these tactics, China-based AI models demonstrate rapid operational adaptation. The US officials said that the Chinese AI model MiniMax redirected exchanges to a new Claude model within 24 hours of release, demonstrating real-time provider monitoring and pre-positioned infrastructure for immediate retargeting.
Chinese AI companies employing industrial-scale distillation against US AI models are able to not only reduce financial expenditures in training a frontier model but also significantly cut AI development timelines.
The US officials have alleged that the Chinese AI company DeepSeek conducted organised distillation campaigns against US frontier AI models, particularly those of Claude, Gemini, GPT, and Grok.
“DeepSeek has been conducting an organized distillation campaign against U.S. AI companies’ frontier AI models since at least late 2024 to generate synthetic training data for its models, including R1, released in early 2025. The company targeted specific knowledge domains to extract proprietary functionality and reasoning capabilities to reduce their compute and research costs. DeepSeek’s publicly quoted training costs of $5.6M are misleading as it does not include the true cost of the data acquired through extensive malicious distillation.”

Similarly, Moonshot AI targeted Claude Fable 5 data to train its Kimi-K3 model and GPT-4o data to train its Kimi-K2 model in mid-2025. Moonshot AI used various models of Claude, GPT, Gemini, Nano Banana, and Grok, distill SFT optimisation, reinforcement learning (RL), software engineering, and math capabilities.

The US CISA has recommended three immediate actions to the US AI companies. It urged the AI companies to implement comprehensive detection and mitigation, deploy targeted response changes, and establish cross-organisation intelligence sharing.
China denies US’s ‘malicious copying of AI tech’ allegations
Hitting back at the US, the Chinese Ministry of Foreign Affairs has asked the US to “refrain from making unfounded accusations or smears” against China.
“China’s AI development is the result of high-level technological self-reliance and strength. We maintain that all parties should strengthen cooperation to promote AI development that is open, inclusive, universally beneficial and oriented toward the common good,” the ministry spokesperson Mao Ning said at a press conference.
Chinese AI models offering features equivalent to those offered by expensive US AI models at way lesser costs
While China denies any wrongdoing the US has accused it of, the unreal price gaps between equivalent US and Chinese AI models raise suspicion.
Pertinently, knowledge distillation trains a student model on the outputs, often reasoning traces, of a stronger teacher model. If done legitimately on one’s own large model or open weights it creates smaller, cheaper and faster models. However, if conducted at industrial scale against a competitor’s APIs without proper permission, distillation allows a lab acquire specialised behaviours without having to pay the full cost of pre-training from scratch or operating equivalent post-training compute.
Chinese labs have, as per US allegations, achieved architectural efficiency, reinforcement-learning techniques for reasoning, like DeepSeek-R1 style, force efficiency from US chip export controls, and aggressive price competition and open-weight releases that recruit global inference providers and fine-tuners.
Consequently, AI models like DeepSeek V3/V4/R1, Alibaba Qwen series, Moonshot Kimi K2/K3, Z.AI GLM, and MiniMax, which score close to or occasionally match mid-to-high US models on coding, math, reasoning and agentic benchmark, while API prices are usually 5 to 7x lower.
For example, while Chinese DeepSeek V4 Flash costs around $0.14/M input and $0.28/M output tokens, the rates of comparable GPT or Claude mid-tiers are far higher. Similarly, while DeepSeek V4 Pro costs around $0.435/$0.87, key flagship US models cost $5/$25–30.
While US labs still lead on the latest flagship models, safety alignments, various enterprise features, and ecosystem lock-in, for many production workloads, the performance-per-dollar gap became significant.
With Chinese models capturing around 30-46% of US enterprise token share on AI API aggregators like OpenRouter at peaks. Result? The enterprises facing high bills moved traffic. In very short time, Chinese AI models emerged as rivals to the US ones.
This forced top US AI companies to cut prices. OpenAI reduced the price of its GPT-5.6 Luna by 80%, to $0.20/M input and $1.20/M output, Terra by 20% though flagship Sol’s high pricing remained unchanged. In fact, Open AI executives have openly compared the new cheaper tiers to the Chinese AI models like Z. AI’s GLM. Anthropic also grappled with similar pressure.
This indicates a classic commoditisation dynamic wherein high fixed costs of frontier pre-training require high inference prices to recoup, while low-cost competitors, be it though open weights, efficiency, subsidies, or alleged free-riding, particularly large-scale distillation campaigns against rival AI models, collapse those prices.
The US allegations suggest that besides efficiency work and open-source distribution, distillation of industrial scale is allowing Chinese AI companies to shorten development timelines and slash compute costs for acquiring specific capabilities that would otherwise need expensive original research and training runs, and thus produce models good enough for most tasks for a fraction of the price high-priced US advanced agentic AI models can perform.
It is notable that Chinese labs have previously published genuine technical contributions and open-source Chinese models have themselves got distilled by others. However, distillation of industrial scale becoming a core strategy, as alleged by the US, has accelerated catch up for Chinese AI models and enabled the price pressure on US AI companies. For US, this is an IP theft and Terms of Use violation, and the US authorities see it as a national security issue, since capabilities were transferred without original safety work.
It appears that Chinese tech entities are countering the strategy of US AI models of producing highly advanced models and charging high prices, by commoditising frontier-level AI. The claims by US officials that Chinese AI companies are doing industrial scale distillation of frontier US AI models lends further credence to the accusations that Chinese AI strategy is nothing but a sophisticated ‘copy the rival’s proprietary functionalities and capabilities for low costs, produce equivalent models, and offer them at way cheaper prices than rivals’ shortcut.


