Ling-3.1-Flash MoE Model Announced with Planned Open-Source Release
Ling-3.1-flash, a Mixture-of-Experts (MoE) language model featuring ~560B total parameters, ~25B active parameters per token, and a 1M-token context window, has been announced. The model is currently offered through a two-week free trial prior to its planned open-source release. Releasing a high-performing 560B-parameter MoE model to the open-source community could substantially elevate the baseline for accessible open AI capabilities. Its sparse activation of only 25B parameters makes inference far more computationally efficient than traditional dense models of comparable total size. The model recorded strong performance across diverse domain benchmarks, achieving 1,673 Elo on GDPVal-AA v2.1 for professional work tasks, 75.16 on FrontierSWE for software engineering, and 65.35 on HealthBench Professional for healthcare applications. Its 1M-token context window enables the processing of extensive codebases and long-form document sets.
## BACKGROUND
A Mixture-of-Experts (MoE) architecture dynamically routes input tokens to specific sub-networks ("experts"), allowing models to scale total parameter count without proportionally increasing inference compute costs. Standard evaluations like GDPVal-AA assess models on economically valuable professional knowledge work, FrontierSWE tests multi-step coding challenge solutions, and HealthBench measures clinical conversation and decision-support abilities.