Generated note

The Timeline Math Behind the Kimi K3 Distillation Accusation

Officials named Fable and banned Nvidia chips in the same breath — but only one of those claims has any supporting detail behind it.

Two claims, bundled together

White House science advisor Michael Kratsios accused Moonshot of two things at once: covertly distilling Anthropic's Fable model, and training on Nvidia GB300 chips banned from export to China. Treasury Secretary Scott Bessent added that US model 'watermarks' show up on Chinese models generally. Neither official shared evidence. Researchers who study distillation say the first claim runs into a wall the moment you check the calendar; the second claim, less discussed, has real infrastructure and precedent behind it.

Where the evidence actually pointsmoderate — based on named researchers' technical reasoning, not on released documents

The distillation-of-Fable claim doesn't survive the timeline: Fable has been public for roughly two weeks, and researchers say that's not enough time to query it at scale, train on the results, and ship a model this strong. The chip-smuggling claim is structurally more plausible — named hardware, a documented black market, a prior indictment — but remains just as unproven, since no one has shown GB300s were actually used to train K3.

What distilling Fable would actually require

01Systematically query Fable

generate training data, sometimes via chain-of-thought prompts — the step Anthropic says it detected in earlier cases

02Fine-tune (SFT) or run reinforcement learning

Lambert: SFT alone is 'becoming less and less impactful' as models approach the frontier

03Large-scale RL training

can require tens of millions of agents; using a frontier lab's API for this would be 'insanely expensive' and slow, per Lambert

04Train and release Kimi K3

happened within roughly two weeks of Fable going public on July 1 — Hancock: 'there's just not even frankly time'

Who has accused whom

accuses of Fable distillation + banned-chip use, no evidence sharedaccuses of Fable …claims to find US model 'watermarks', unspecifiedclaims to find US…earlier 2026 accusation: IP-tracked distillation queriesearlier 2026 accu…same earlier accusationsame earlier accu…same earlier accusationsame earlier accu…alleged distillation target — unconfirmed for K3 specificallyalleged distillat…Michael KratsiosScott BessentAnthropicMoonshot / Kimi K3DeepSeekMiniMaxFable (Anthropic model)

Anthropic's earlier accusation named a detection method (IP-tracked query patterns). Kratsios and Bessent have not disclosed theirs.

Distillation claim vs. chip-smuggling claim

Distillation (Fable → K3)

  • Fable public since July 1; K3 shipped ~2 weeks later — 'not even time,' per Hancock
  • Would need expensive, slow large-scale RL run against Anthropic's API — 'insanely expensive,' per Lambert
  • No specifics shared on Bessent's alleged 'watermarks' or Kratsios's sourcing
  • Moonshot declined to answer questions about its training process

Banned-chip smuggling

  • Specific hardware named: Grace Blackwell 300s, plus GB300-equipped servers in Thailand
  • Black market for banned Nvidia chips is documented (Bresnick, Georgetown CSET)
  • Supermicro's founder was indicted in May for smuggling advanced chips into China
  • No federal know-your-customer rule for data centers has been enacted since the 2024 proposal
Inside: why SFT and reinforcement learning have different time costs

Distillation covers two different processes. Supervised fine-tuning (SFT) trains a model on prompt-response pairs pulled from the target — Lambert calls this where 'the model picks up its manners,' and it's the mechanism behind models that mistakenly claim to be Claude. But SFT's payoff is shrinking as target models get more complex. Matching Fable-level capability would more likely require reinforcement learning, where an agent of the larger model grades the smaller model's outputs — a process that needs tens of millions of agents and heavy infrastructure, and would be slow and costly to run against a commercial API. That's the technical basis for Hancock and Lambert's skepticism: cheap SFT could plausibly happen fast; the RL work that would actually explain K3's strength could not.

The context officials left out

context

Elon Musk testified that his company SpaceXAI distilled OpenAI models to build Grok, calling the practice common across the industry — not unique to Chinese labs.

context

One of Moonshot's founders is a CMU PhD; Hancock warns against assuming Chinese teams are 'just riding coattails' rather than doing original work.

unresolved

K3 may still owe something to older Anthropic models — Anthropic's earlier, evidence-backed accusation covered prior distillation — even if it did not distill Fable specifically within two weeks.

Sources