Engineer Trains 102M Recursive BitNet-v2 Model with 64K Context
An independent engineer trained a 102M-parameter recursive BitNet-v2 model from scratch with a 64K context window using less than 5 billion tokens. The model combines ternary weights, shared transformer layers, and hashed n-gram embeddings on a tiny budget. This project demonstrates how experimental and cost-effective combinations of efficient LLM architectures can yield functional models with long context windows. It highlights the potential of resource-constrained training methodologies in the open-source AI community. The decoder-only transformer features 102.28M unique parameters, utilizes six transformer blocks run twice with shared weights, and incorporates hashed 2-, 3-, and 4-gram embeddings. The model was trained in two stages using NVIDIA B300 GPUs and achieved an unweighted zero-shot benchmark average of 40.59%.
## BACKGROUND
Ternary Weight Networks quantize model weights to {-1, 0, +1}, drastically cutting computational cost and memory use. N-gram embeddings explicitly incorporate multi-token patterns into the model, while recursive or shared transformer layers reuse parameters across depth to increase effective capacity without adding unique parameters.