Key Lessons and Cost Trade-offs from Building the Apex-2 Language Model
An independent developer shared practical takeaways from pre-training the Apex-2 model on a $2,000 budget, highlighting the effective use of DiLoCo distributed training across two distinct GH200 GPU nodes. By merging model weights every 350 steps, the setup achieved a 1.9x training speedup at approximately 40% Model FLOPs Utilization (MFU) per GPU compared to single-GPU training. This experiment demonstrates how distributed algorithms like DiLoCo can make multi-node LLM training accessible without requiring expensive ultra-high-bandwidth interconnects like NVLink between nodes. It offers valuable empirical data for resource-constrained researchers striving to optimize dataset deduplication and hardware compute efficiency. Cross-snapshot deduplication using MinHash removed 57% of the FineWeb-Edu sample and 34% of DCLM, though the author noted this deduplication step is a budget trade-off rather than an absolute quality win. Due to computational and monetary constraints, the final dataset size was scaled down from an initial target of 1 Trillion tokens to 80 Billion tokens for a Mixtral-style Mixture-of-Experts (MoE) architecture.
## BACKGROUND
Distributed Low-Communication (DiLoCo) training is an optimization algorithm that allows training large language models across poorly connected compute clusters by communicating updates only after multiple local steps. Model FLOPs Utilization (MFU) measures the ratio of executed floating-point operations to the theoretical maximum peak performance of the GPU. FineWeb-Edu is a curated web dataset released by Hugging Face, filtered using LLM classifiers to retain high-quality educational text.