Recent allegations have surfaced suggesting that Moonshot, the Chinese company behind the Kimi K3 large language model (LLM), developed its model by illicitly copying Anthropic’s Fable LLM and utilizing unauthorized hardware. White House science advisor Michael Kratsios accused Moonshot of engaging in large-scale industrial distillation to appropriate proprietary U.S. technology, a claim that has sparked significant debate within the AI community.
Kratsios’s assertions align with comments from Treasury Secretary Scott Bessent, who indicated that U.S. LLM watermarks have been detected in several Chinese models, labeling this as unacceptable. However, specifics regarding these watermarks and the evidence supporting these claims remain undisclosed.
Despite these serious allegations, several AI experts express skepticism about the feasibility of such distillation methods leading to the rapid development of a model as advanced as Kimi K3. Braden Hancock, a researcher at the Laude Institute and co-founder of Snorkel AI, pointed out the impracticality of achieving such progress solely through distillation in a short timeframe, especially considering Fable’s public availability since July 1st.
Similarly, Nathan Lambert, an AI researcher at the Allen Institute for AI, noted in a recent podcast that distillation’s impact has diminished as Chinese models approach the frontier. He emphasized that if distillation were sufficient, other models would have easily caught up to Kimi K3’s capabilities through supervised fine-tuning alone, which has not been observed.
Distillation involves systematically querying a target model to generate data for post-training, often requiring the model to articulate its problem-solving processes. This data is then used to train a new model through supervised fine-tuning (SFT). However, Lambert suggests that as models become more complex, SFT’s benefits are waning, and achieving Fable-like capabilities would likely necessitate reinforcement learning techniques. These advanced methods demand substantial infrastructure, including large-scale reinforcement learning runs involving tens of millions of agents, making the process exceedingly resource-intensive and costly.
In light of these insights, the AI community remains divided on the plausibility of Moonshot’s alleged methods. While the accusations highlight concerns over intellectual property and technological ethics, the technical challenges associated with distillation and reinforcement learning suggest that Kimi K3’s development may not be as straightforward as claimed. This ongoing discourse underscores the need for transparency and collaboration in AI development to address security concerns without stifling innovation.