ByteDance Reportedly Planning to Train a 5-Trillion-Parameter AI Model
ByteDance's Seed Foundation team is reportedly discussing plans to train a massive AI model with over 5 trillion parameters, which would make it the largest known model in China. The project is led by Xiang Liang and Shen Ke, key figures in ByteDance's AI and large model development. This move represents a major escalation in the AI race within China, aiming to leapfrog competitors like Alibaba and Moonshot AI by pushing the limits of model scale. ByteDance founder Zhang Yiming has strategically backed this effort, emphasizing the pursuit of intelligence limits over short-term gains from model distillation. Some insiders view the massive scale jump as a gamble, but Zhang Yiming has reassured the team, stating he accepts temporary lag behind competitors to focus on long-term breakthroughs. He specifically highlighted programming as a key direction and expressed opposition to knowledge distillation, arguing it merely replicates existing models like Claude.
## BACKGROUND
Large language models rely on parameters—variables that the model learns during training—to process and generate text, where larger parameter counts generally correlate with higher intelligence. Knowledge distillation is a technique where a smaller, more efficient "student" model is trained to reproduce the behavior of a larger "teacher" model, which can speed up deployment but may limit original capabilities.