Researchers Use Machine Unlearning to Remove CCP Censorship from Qwen LLMs
Researchers at Hirundo released open-weight "Westernized" versions of Qwen3.6-35B-A3B and Qwen3.5-4B, applying targeted machine unlearning to drop CCP-aligned political censorship and propaganda responses from 89.8% to 2.8%. The process preserved general model capabilities, with benchmark scores moving by an average of only 0.72 points. This work demonstrates a precise method for removing embedded political biases and state censorship from open-weights LLMs without destroying general reasoning capabilities or basic safety guardrails. It offers the open-source AI community an effective alternative to crude activation-projection methods like abliteration when re-aligning base models for different geopolitical contexts. Unlike abliteration, which strips a model's general ability to refuse prompts, Hirundo's approach trains a LoRA adapter with a behavioral-unlearning objective while using a retain set to preserve overall benchmark performance and safety features. Rather than substituting an opposing political dogma, the unlearned model presents neutral, multi-perspective viewpoints on sensitive topics such as Taiwan's sovereignty.
## BACKGROUND
Machine unlearning in large language models refers to techniques designed to selectively erase specific knowledge, behaviors, or biases from a model's trained weights after initial training. Open-weight Chinese LLMs like Alibaba's Qwen family are technically state-of-the-art but incorporate strict censorship alignment regarding Chinese political topics, which persists in model weights even when system prompts are altered.