Back to News Hub
🟢TechCrunch AI
July 23, 2026
General AI

Experts say exploiting Anthropic's Fable isn't how Kimi K3 got so good

Overview

"I don't think you get a model this strong and this quickly on the heels of Fable doing strictly distillation," one expert told TechCrunch. White House science advisor Michael Kratsios said that Moonshot, the Chinese company behind Kimi K3, the largest available open-weight LLM, built its model by copying Anthropic's Fable LLM while using chips that aren't cleared for export to China. "Large-scale, covert industrial distillation aimed at stealing proprietary U.

Key Takeaways

  • technology and undermining American research is unacceptable," Kratsios wrote , amid reported discussions about banning Chinese open-weight models that have roiled the AI sector.

    Moonshot did not respond to questions about its training process, and Kratsios did not share more details about the sources of his allegations.

  • "I don't think you get a model this strong and this quickly on the heels of Fable doing strictly distillation," Braden Hancock, a researcher at the Laude Institute and co-founder of Snorkel AI, told TechCrunch.
  • " Performing distillation requires a lab to systematically query its target model in order to generate data that can be used for post-training.

    Sometimes this explicitly involves asking the model to articulate its chain-of-thought to understand how it solves problems.

  • To distill Fable-like capabilities would likely require reinforcement learning techniques.

    In many cases, that means having an agent of the larger model grade the smaller model's responses, and adjusting based on the grade.

  • Those queries were "distinct from normal usage patterns, reflecting deliberate capability extraction rather than legitimate use.

technology and undermining American research is unacceptable," Kratsios wrote , amid reported discussions about banning Chinese open-weight models that have roiled the AI sector. Moonshot did not respond to questions about its training process, and Kratsios did not share more details about the sources of his allegations. Kratsios' tweet echoed comments from Treasury Secretary Scott Bessent that "we are finding watermarks of our U.

large language models on many of the Chinese models, and that that's unacceptable. " It's not clear what those watermarks consist of, and the Treasury Department did not respond to a query. However, experts are skeptical that distillation - the process of querying an LLM to determine its inner workings and copy its capabilities - is responsible for the advanced capabilities that Kimi K3 displays.

You can't distill that much data, train a model, and release it in two weeks. " "I've been of the opinion that distillation is becoming less and less impactful over time as the Chinese models get closer to the frontier and the training regime shifts to [reinforcement learning]," Nathan Lambert, an AI researcher at the Allen Institute for AI, said in a podcast released yesterday. "[I]f it were the case, everyone would be easily able to catch up to a GLM or to a K3 by using its data for distillation.

For more details please read the original article at TechCrunch AI.

Continue Learning

Originally published by TechCrunch AI
Read the original

Comments

Sign in to join the conversation