DeepSeek Censorship Doesn't Transfer via Distillation

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- DeepSeek V4 Flash scored 45.45 points more censored on China-sensitive questions than on structurally identical non-China controls across 152 matched prompt pairs, evaluated by judges from xAI, Google, OpenAI, and Anthropic.
- The distilled GPT-OSS-120B displayed no statistically significant difference in political-topic behavior from its untouched base model, sitting at +3.94, +3.70, and +2.58 censorship-gap points across three arms versus the teacher's +32.02.
- Self-distillation matched DeepSeek distillation on every seed, with the self-taught arm reaching 83.61% on FinanceReasoning at an 8,000-token budget — at $0.00025939 per query, 62 times lower cost than Inkling and 160 times lower than Kimi K3.
- At a 100,000-token budget the large models win on raw accuracy (Kimi K3 at 89.92%, Inkling at 88.24%, DeepSeek V4 Flash at 85.71%), but at the production-realistic 8k budget Kimi K3 falls to 81.93% and Inkling to 65.13%, while the 120B completes 98.7% of problems and doesn't move.
- LineageEval, the open-source evaluation apparatus (304 prompts in 152 matched pairs, judge rubric, and evaluation code), was published alongside the models themselves to let practitioners reproduce the censorship-transfer audit.
- A 20B open-weight variant required expert-layer adaptation after attention-only tuning underfit at that scale, reaching 74.79% on FinanceReasoning versus 64.71% for its base, served in 42 GB of weights.
Why it matters: Practitioners distilling from Chinese frontier models for cost or performance reasons now have empirical evidence that political censorship does not carry over to the student — and that self-distillation matches DeepSeek distillation at 12.5% fewer output tokens, with no external model appearing in the training signal. For finance-adjacent deployment, the 120B ships from a single H100 or A100 at roughly 63 GB of native MXFP4 weights plus an 80 MB attention-only adapter.



