Anthropic publicly accused three Chinese AI labs — DeepSeek, Moonshot AI, and MiniMax — of running industrial-scale distillation campaigns against Claude, generating over 16 million exchanges via fraudulent accounts. The analysis breaks down the actual scale and likely impact: DeepSeek's 150K exchanges are negligible for training purposes, while MiniMax's 13M+ exchanges represent a more substantial dataset. The author argues distillation's impact is real but overstated — it helps with SFT-style post-training but cannot replace on-policy RL compute, which is now central to frontier model development. Chinese labs likely innovate heavily on distillation efficiency due to GPU access restrictions, but this alone won't close the frontier gap. The piece also contextualizes the geopolitical dimension, noting that restricting API-based distillation is far harder than restricting GPU exports, and that Anthropic faces a fundamental tension between offering a competitive API and protecting its model capabilities.

11m read timeFrom interconnects.ai
Post cover image
23 Impressions