~/MACHINE LEAR/ml-developer-trains-3-87b-moe-model-from-scratch-on-86-5b

ML Developer Trains 3.87B MoE Model from Scratch on 86.5B Tokens

An open-source developer trained Apex-2, a 3.87 billion parameter Mixture-of-Experts (MoE) language model with 1.45 billion active parameters per token, entirely from scratch using only 86.5 billion pretraining tokens. Utilizing GH200 GPUs and the DiLoCo distributed training algorithm, the model achieved coding benchmark performance competitive with models trained on vastly larger datasets, such as Qwen2.5-1.5B. This project demonstrates that independent researchers can pretrain high-performing custom MoE architectures on constrained compute budgets and limited token volumes. It highlights the practical efficiency of sparse MoE designs and communication-efficient distributed training algorithms like DiLoCo in reducing resource barriers for LLM development. Apex-2 consists of 32 layers with 16 experts per layer (top-4 routing), Grouped-Query Attention (GQA), and a 4,096 context length using the Qwen3 tokenizer. While Supervised Fine-Tuning (SFT) yielded strong code benchmark scores (HumanEval 43.9, HumanEval+ 41.5), applying Direct Preference Optimization (DPO) degraded math and coding performance, prompting the author to keep the SFT checkpoint.

## BACKGROUND

A Mixture-of-Experts (MoE) architecture routes input tokens to a subset of specialized subnetworks (experts), expanding total model parameters while keeping per-token computational costs low. DiLoCo (Distributed Low-Communication) is a training algorithm that dramatically reduces cross-device communication overhead, allowing efficient LLM training across poorly connected GPU clusters. Direct Preference Optimization (DPO) is an alignment technique designed to steer model outputs toward human preferences without training a separate reward model.

## REFERENCES

## KEYWORDS

#machine-learning#llm#mixture-of-experts#model-training#open-source-ai

$ subscribe --daily