~/LLM FINE TUN/developer-shares-post-training-challenges-distilling-chain-of-thought-into-aliceai-80b

Developer Shares Post-Training Challenges Distilling Chain-of-Thought into AliceAI-80B

An open-source developer shared progress on locally fine-tuning Yandex's AliceAI-80B-A3B model using three 32GB Nvidia V100 GPUs. After an initial round of supervised fine-tuning (SFT) on 5 million tokens resulted in severe underfitting for general conversation, the developer expanded synthetic data generation using their open-source tool, sftmill. This practical experiment demonstrates both the feasibility and token-breadth requirements of post-training modern Mixture-of-Experts (MoE) models on older enterprise hardware. It highlights how accessible off-policy distillation engines enable individual developers to teach open models specialized reasoning capabilities. The first checkpoint successfully learned basic chain-of-thought reasoning but produced garbled outputs on ambiguous queries due to a narrow dataset focused primarily on coding. To fix this, synthetic data generation throughput was tripled to 240 tokens per second across multiple parallel instances, preparing 5 million broader general-instruction tokens for a second training pass.

## BACKGROUND

Yandex's AliceAI-Foundation-80B-A3B-Base is an open-weight base model utilizing a Mixture-of-Experts (MoE) architecture with 80 billion total parameters, of which 3 billion are active per token. Model distillation and Supervised Fine-Tuning (SFT) allow developers to adapt base models by training them on synthetic outputs generated by stronger teacher models to impart instruction-following and reasoning abilities.

## REFERENCES

## KEYWORDS

#LLM Fine-Tuning#Model Distillation#LocalLLaMA#Open Source AI

$ subscribe --daily

Developer Shares Post-Training Challenges Distilling Chain-of-Thought into AliceAI-80B | Daily News