Fine-Tuning Progress Update on Yandex's AliceAI 80B-A3B Model
Reddit user jjusko20 shared a progress update on fine-tuning Yandex's AliceAI 80B-A3B foundation model, reaching approximately 40% completion of the initial training run. The update includes a training loss curve logged from the final micro-step of each iteration along with a public live-stream link for the training run. Community fine-tuning efforts on high-capacity open-weight models allow researchers and independent developers to evaluate practical post-training workflows on Mixture-of-Experts (MoE) architectures. Sharing loss curves and live training streams provides transparent empirical reference points for engineers tuning large sparse models. surrounding variations in the loss curve were explained by sampling metrics specifically from the final micro-step of each iteration, with setup assistance from Gemini 3.8 Flash High. The training target is Yandex's Apache 2.0 licensed AliceAI-Foundation-80B-A3B base model.
## BACKGROUND
Yandex open-sourced the AliceAI-Foundation-80B-A3B base language model under the Apache 2.0 license, featuring a hybrid Mixture-of-Experts (MoE) architecture. The model contains 80 billion total parameters, but only activates 3 billion parameters per token during inference while supporting context lengths up to 262,144 tokens. MoE architectures route incoming tokens dynamically to specialized sub-networks, drastically lowering active compute requirements compared to standard dense models.