Skip to main content
Back to News Hub
🟢TechCrunch AI
July 23, 2026
Claude

Experts say exploiting Anthropic's Fable isn't how Kimi K3 got so good

Overview

"I don't think you get a model this strong and this quickly on the heels of Fable doing strictly distillation," one expert told TechCrunch. White House science advisor Michael Kratsios said that Moonshot, the Chinese company behind Kimi K3, the largest available open-weight LLM, built its model by copying Anthropic's Fable LLM while using chips that aren't cleared for export to China. "Large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology and undermining American research is unacceptable," Kratsios wrote , amid reported discussions about banning Chinese open-weight models that have roiled the AI sector.

Key Takeaways

  • Moonshot did not respond to questions about its training process, and Kratsios did not share more details about the sources of his allegations.

    Kratsios' tweet echoed comments from Treasury Secretary Scott Bessent that "we are finding watermarks of our U.S. large language models on many of the Chinese models, and that that's unacceptable."

  • Fable's only been publicly available since July 1st.

    You can't distill that much data, train a model, and release it in two weeks."

  • Sometimes this explicitly involves asking the model to articulate its chain-of-thought to understand how it solves problems.

    Other times, the prompts and responses from a model are used to train a new model in a process called supervised fine-tuning, or SFT.

  • In many cases, that means having an agent of the larger model grade the smaller model's responses, and adjusting based on the grade.

    The more advanced techniques also require more significant infrastructure.

  • Those queries were "distinct from normal usage patterns, reflecting deliberate capability extraction rather than legitimate use."

Moonshot did not respond to questions about its training process, and Kratsios did not share more details about the sources of his allegations. Kratsios' tweet echoed comments from Treasury Secretary Scott Bessent that "we are finding watermarks of our U.S. large language models on many of the Chinese models, and that that's unacceptable." It's not clear what those watermarks consist of, and the Treasury Department did not respond to a query.

However, experts are skeptical that distillation - the process of querying an LLM to determine its inner workings and copy its capabilities - is responsible for the advanced capabilities that Kimi K3 displays. "I don't think you get a model this strong and this quickly on the heels of Fable doing strictly distillation," Braden Hancock, a researcher at the Laude Institute and co-founder of Snorkel AI, told TechCrunch. "There's just not even frankly time, right?

Fable's only been publicly available since July 1st. You can't distill that much data, train a model, and release it in two weeks." "I've been of the opinion that distillation is becoming less and less impactful over time as the Chinese models get closer to the frontier and the training regime shifts to [reinforcement learning]," Nathan Lambert, an AI researcher at the Allen Institute for AI, said in a podcast released yesterday.

For more details please read the original article at TechCrunch AI.

Continue Learning

Comments

Comments appear only after moderation. Your email identifies your submission to the moderator and is never displayed here.

No approved comments yet.

Originally published by TechCrunch AI
Read the original