Once a niche topic for AI researchers, “distillation”—the practice of training smaller AI models using outputs from more advanced ones—has rapidly become a central point of debate in Silicon Valley and Washington D.C. Triggered by Chinese firm Moonshot AI’s competitive Kimi K3 model, which allegedly distilled Anthropic’s Fable, this technique is now raising national security and intellectual property theft concerns. While U.S. policymakers weigh restrictions, tech giants like Nvidia and Microsoft are advocating against premature curbs, arguing distillation is a legitimate innovation tool crucial for competition.
Once confined to advanced AI research circles, the concept of 'distillation' has exploded into mainstream tech discourse, sparking intense debate from Silicon Valley boardrooms to Washington D.C. This shift comes as the quick ascent of Chinese AI firm Moonshot AI's Kimi K3 model forces tech innovators and lawmakers to confront a powerful, yet controversial, technique.
Earlier this year, Google AI lead Jeff Dean offered an early glimpse into distillation on a podcast, explaining how Google utilized these techniques to enhance system performance without relying solely on large image recognition models. Dean stated in February, "Through distillation, which is a key technique for making the smaller models more capable, you have to have the frontier model in order to then distill it into your smaller model."

Just five months later, distillation has emerged as a hot-button issue, fueling concerns that it could pose a national security threat and allow China to rapidly close the gap in the critical AI race. These anxieties escalated following the release of Moonshot AI's Kimi K3, which users quickly found to be remarkably competitive with leading commercial AI models from companies like Anthropic and OpenAI.
A key differentiator is that Moonshot AI and other Chinese labs are increasingly offering 'open-weight' models, enabling users to download, customize, and deploy the technology freely. This stands in contrast to U.S. giants that typically sell access to proprietary, closed models.
Some U.S. government officials allege that Moonshot's rapid advancement is a result of distillation, characterizing it as intellectual property theft, particularly pointing to the incorporation of Anthropic's advanced Fable model. White House advisor Michael Kratsios publicly stated on X, "We have information that Moonshot AI distilled Anthropic's Fable for the development of its K3 model. To do this they developed a sophisticated internal platform to conduct large scale distillation against U.S. models, allowing them to quickly switch between multiple methods of access to avoid detection."
At its core, distillation involves using the outputs of a highly advanced AI model or chatbot to train a simpler, often smaller, model. This practice is controversial because it can allow developers to create competitive AI offerings without investing the immense capital and resources required to develop the original, cutting-edge training technology.
Pukar Hamal, founder of AI security firm SecurityPal, likened it to academic dishonesty: "It's almost like someone went to the lectures, read the textbook, and did all the hard work of doing the homework. Then some other student is like, 'Hey, I didn't do that. Can I just copy your work?'"
In a striking display of industry unity, major tech players including Nvidia, Microsoft, Meta, and Palantir, alongside over 20 other companies, issued a joint letter urging policymakers to avoid "premature restrictions" on open-weight AI models. They argued that such restrictions would "stifle competition or drive innovation overseas." The letter explicitly defended distillation, stating, "Distillation, or the practice of using one model's outputs to help train or improve another, is a widely used technique for model improvement, evolution, and validation."
Complicating the China problem
The rise of distillation presents a complex challenge for U.S. policymakers, who have long grappled with Chinese technological advancements, intellectual property concerns, and national security implications. Colin Shea-Blymyer, a research fellow at Georgetown's Center for Security and Emerging Technology, noted the U.S. government is actively trying to define its stance.
Shea-Blymyer suggested the government might argue that Chinese and Russian companies have gained an "unfair advantage" by using outputs from "hardworking American models" to boost their own performance. However, Aaron Levie, CEO of Box and a signatory of the industry letter, emphasized the necessity for U.S. companies to access the best available technology, regardless of its origin, to remain competitive. He believes that greater innovation globally will ultimately lead to more affordable and efficient AI.
While much of the current debate centers on Chinese open-weight models like Kimi K3, distillation is a technique widely adopted across the industry. Shashi Bellamkonda, research director at Info-Tech Research Group, highlighted that Nvidia, for instance, employed distillation in training its Llama Nemotron series of models, as detailed in their research papers. "It is a legitimate and a very valuable technique to train a smaller, cheaper model on outputs of a larger model, and is practiced all the time," Bellamkonda explained.

Anthropic, however, holds a contrasting view, particularly as its proprietary models are targeted. In February, the company reported that its Claude capabilities were being distilled on an "industrial scale" by Chinese firms like DeepSeek, Moonshot, and MiniMax, involving approximately 24,000 fake accounts and generating 16 million exchanges. Valued at nearly $1 trillion and eyeing a potential IPO, Anthropic considers stopping illicit distillation a matter of national security. They stated that preventing state and non-state actors from leveraging AI for harmful activities like bioweapons development or cyberattacks necessitates "rapid, coordinated action among industry players, policymakers, and the global AI community."
Both OpenAI and Anthropic have implemented terms of service banning distillation, viewing unauthorized use of their larger models as potential IP theft. Yet, with AI development costs soaring, companies are constantly seeking efficiencies.
Hamal of SecurityPal indicated his company would consider using Chinese open-weight models like Kimi K3, provided there are no "nefarious backdoors in the code," citing significant cost savings. He questioned, "But hosting it on our own infrastructure after we've done an assessment, why not?"
A critical challenge for companies like Anthropic and OpenAI in their IP theft arguments is their own history of utilizing various content sources to build their models, which has led to copyright lawsuits against them. Max Pritt, an attorney for Boies Schiller Flexner representing authors in copyright litigation against AI firms, highlighted this inconsistency: "The administration, at least publicly, has focused its efforts on the protection of technology companies' intellectual property, while remaining silent in large part about creators and individuals' intellectual property that was used without authorization."
