~/MACHINE LEAR/engineer-trains-102m-recursive-bitnet-v2-model-with-64k-context

Engineer Trains 102M Recursive BitNet-v2 Model with 64K Context

An independent engineer trained a 102M-parameter recursive BitNet-v2 model from scratch with a 64K context window using less than 5 billion tokens. The model combines ternary weights, shared transformer layers, and hashed n-gram embeddings on a tiny budget. This project demonstrates how experimental and cost-effective combinations of efficient LLM architectures can yield functional models with long context windows. It highlights the potential of resource-constrained training methodologies in the open-source AI community. The decoder-only transformer features 102.28M unique parameters, utilizes six transformer blocks run twice with shared weights, and incorporates hashed 2-, 3-, and 4-gram embeddings. The model was trained in two stages using NVIDIA B300 GPUs and achieved an unweighted zero-shot benchmark average of 40.59%.

## BACKGROUND

Ternary Weight Networks quantize model weights to {-1, 0, +1}, drastically cutting computational cost and memory use. N-gram embeddings explicitly incorporate multi-token patterns into the model, while recursive or shared transformer layers reuse parameters across depth to increase effective capacity without adding unique parameters.

## REFERENCES

## KEYWORDS

#Machine Learning#LLM Architecture#BitNet#Model Training#Efficient AI

$ subscribe --daily

Engineer Trains 102M Recursive BitNet-v2 Model with 64K Context | Daily News