
Chinese military researchers have successfully trained domestic AI systems to improve China's defense capabilities using outputs from top US AI models created by OpenAI and Anthropic, according to a Reuters analysis of more than 80 Chinese academic publications and patents. Despite Washington's efforts to limit Beijing's access to cutting-edge chips and other strategic technologies, the findings provide insights into how Chinese military and security-related institutions are using cutting-edge US AI models as a shortcut to developing specialized systems. Researchers connected to the People's Liberation Army and other military organizations frequently employ model distillation, a method that uses outputs of strong AI systems to train smaller, more specialized models that can be implemented locally without massive processing capacity. As per Sunny Cheung, a Jamestown fellow who analyzed over 60 of the papers, Chinese military scientists are systematically capturing the reasoning steps of Western models to adapt them for surveillance, cyberwarfare, and tactical decision-making, noting that "teaching a model the right answer is one thing but teaching it the reasoning behind the answer is much harder."
Chinese defense organizations view cutting-edge American AI models as a source of technological knowledge and a means of catching up to American competitors, as reported by Reuters. According to the analysis, Chinese military scientists are methodically documenting the reasoning steps of Western models to adapt them for surveillance, cyberwarfare, and tactical decision-making. PLA Unit 96941, a military intelligence and cyber-warfare unit in Beijing, revealed using OpenAI's GPT-3.5 to parse classified military source code in a study published last year, where researchers said third-party models were unsuitable for handling classified information. To overcome this limitation, they used GPT-3.5 to summarize software code and trained a domestic model on those summaries to run entirely within Chinese military networks. Researchers at the North University of China, which has strong ties to the nation's weapons industry, created artificial training data for a text classification model for social media monitoring using Anthropic's Claude 3 Haiku. Anthropic told Reuters it does not provide commercial access to Claude in China or to Beijing-controlled firms and uses monitoring systems to detect violations of its policies. Researchers at the PLA's National University of Defense Technology used distillation to compress an image-processing model for deployment on unmanned aerial vehicles, allowing drones to process live video and support navigation and targeting even when communications were disrupted. Meanwhile, scientists at the Academy of Military Sciences developed a target-recognition system for simulated maritime operations involving drones, ships and unmanned submarines.
The subject has become a key flashpoint ahead of US-China negotiations on AI governance and safety, with some Chinese organizations allegedly using distillation to extract capabilities from American AI models, according to US officials. China has rejected the charges, claiming that Washington is seeking AI 'hegemonism' and that identical actions have been taken by American companies. Last week, Moonshot, an AI startup, refuted claims made by the Trump administration that their Kimi K3 model was developed by distillation, claiming it was powered by exclusive advancements. According to Anthropic, it employs monitoring mechanisms to identify regulatory infractions and does not grant commercial access to Beijing-controlled companies or to Claude in China. The dispute centers on unauthorized extraction, not distillation itself, a widely used industry practice. The company added that distilled models may lose the original systems' safety safeguards, potentially allowing sensitive capabilities to be transferred to models beyond its control. The White House, Pentagon, Chinese Foreign Ministry, PLA and OpenAI did not respond to Reuters' requests for comment.
In a 2024 paper, the PLA's National University of Defense Technology explained how to use distillation to reduce the size of an image-processing model for use on unmanned aerial vehicles, enabling drones to analyze live video and support real-time navigation and targeting decisions even during communication disruptions. According to a paper released earlier this year, researchers at China's Academy of Military Sciences employed distillation to run a target-recognition model on tactical gear during simulated maritime operations, including drones, ships, and unmanned submarines. With Washington's export restrictions on high-end CPUs limiting access to advanced computing capabilities, China has adopted distillation as a means of competing with the United States in frontier AI. Central and local governments have promoted "model lightweighting" and edge computing, directing subsidies and research funding toward technologies that enable AI models to run on drones, satellites and other devices with limited processing power. Outside direct military use, researchers at the North University of China, which has close links to the country's defence industry, used Anthropic's Claude 3 Haiku to generate synthetic training data for social media monitoring and content moderation.
Military experts are examining distillation as a potential security problem as Chinese AI models catch up to their American counterparts. The possibility of 'data-free distillation,' a technique for reverse-engineering a model's capabilities without direct access to its fundamental parameters, was discussed in a paper published in January by researchers at the Army Engineering University. They suggested defense measures intended to conceal the concealed logical information revealed in a model's public outputs. According to Trevor Koverko, co-founder of Sapien, distilled models are still not as good as their instructor models, as they only inherit specific characteristics and cannot fully mimic the wide-ranging intelligence of frontier systems. He noted that distilled models remain less capable than their teacher models, best understood as transferring selected capabilities into a cheaper, locally controlled system, not achieving independence from frontier AI. Chinese military researchers are also studying distillation as a potential security threat as domestic AI systems become more capable. Experts noted that while distillation can transfer specific capabilities into smaller, locally controlled systems, it cannot fully reproduce the broad intelligence of frontier AI models.