
The Timeline Math Behind the Kimi K3 Distillation Accusation
Officials named Fable and banned Nvidia chips in the same breath — but only one of those claims has any supporting detail behind it.
Two claims, bundled together
White House science advisor Michael Kratsios accused Moonshot of two things at once: covertly distilling Anthropic's Fable model, and training on Nvidia GB300 chips banned from export to China. Treasury Secretary Scott Bessent added that US model 'watermarks' show up on Chinese models generally. Neither official shared evidence. Researchers who study distillation say the first claim runs into a wall the moment you check the calendar; the second claim, less discussed, has real infrastructure and precedent behind it.
The distillation-of-Fable claim doesn't survive the timeline: Fable has been public for roughly two weeks, and researchers say that's not enough time to query it at scale, train on the results, and ship a model this strong. The chip-smuggling claim is structurally more plausible — named hardware, a documented black market, a prior indictment — but remains just as unproven, since no one has shown GB300s were actually used to train K3.
What distilling Fable would actually require
generate training data, sometimes via chain-of-thought prompts — the step Anthropic says it detected in earlier cases
Lambert: SFT alone is 'becoming less and less impactful' as models approach the frontier
can require tens of millions of agents; using a frontier lab's API for this would be 'insanely expensive' and slow, per Lambert
happened within roughly two weeks of Fable going public on July 1 — Hancock: 'there's just not even frankly time'
Who has accused whom
Anthropic's earlier accusation named a detection method (IP-tracked query patterns). Kratsios and Bessent have not disclosed theirs.
Distillation claim vs. chip-smuggling claim
Distillation (Fable → K3)
- Fable public since July 1; K3 shipped ~2 weeks later — 'not even time,' per Hancock
- Would need expensive, slow large-scale RL run against Anthropic's API — 'insanely expensive,' per Lambert
- No specifics shared on Bessent's alleged 'watermarks' or Kratsios's sourcing
- Moonshot declined to answer questions about its training process
Banned-chip smuggling
- Specific hardware named: Grace Blackwell 300s, plus GB300-equipped servers in Thailand
- Black market for banned Nvidia chips is documented (Bresnick, Georgetown CSET)
- Supermicro's founder was indicted in May for smuggling advanced chips into China
- No federal know-your-customer rule for data centers has been enacted since the 2024 proposal
Inside: why SFT and reinforcement learning have different time costs
Distillation covers two different processes. Supervised fine-tuning (SFT) trains a model on prompt-response pairs pulled from the target — Lambert calls this where 'the model picks up its manners,' and it's the mechanism behind models that mistakenly claim to be Claude. But SFT's payoff is shrinking as target models get more complex. Matching Fable-level capability would more likely require reinforcement learning, where an agent of the larger model grades the smaller model's outputs — a process that needs tens of millions of agents and heavy infrastructure, and would be slow and costly to run against a commercial API. That's the technical basis for Hancock and Lambert's skepticism: cheap SFT could plausibly happen fast; the RL work that would actually explain K3's strength could not.
The context officials left out
Elon Musk testified that his company SpaceXAI distilled OpenAI models to build Grok, calling the practice common across the industry — not unique to Chinese labs.
One of Moonshot's founders is a CMU PhD; Hancock warns against assuming Chinese teams are 'just riding coattails' rather than doing original work.
K3 may still owe something to older Anthropic models — Anthropic's earlier, evidence-backed accusation covered prior distillation — even if it did not distill Fable specifically within two weeks.