~/MODEL COMPRE/a-developer-compresses-a-local-story-generating-llm-into-a-24kb-html

A Developer Compresses a Local Story-Generating LLM into a 24KB HTML File

A developer created a completely self-contained 24KB HTML file that houses both a micro LLM and its JavaScript inference engine, enabling local story generation in a browser at over 60 tokens per second on smartphones. The project compressed an existing 81KB FP32 MacroStory model through adaptive quantization, Quantization-Aware Training (QAT), and aggressive HTML/JS minification. While built as a technical experiment, this project demonstrates the extreme limits of model compression and zero-dependency edge AI. It highlights how custom optimization can deliver private, instant AI inference directly within standard web browsers without relying on external servers or APIs. Unlike standard LLMs where tensor weights dominate memory, embeddings account for roughly 60% of tiny models, making them uniquely sensitive to compression overhead. Popular low-bit techniques like LittleBit or GGUF K-quants proved unsuitable because their metadata overhead exceeded the saved space, prompting the creator to use a tailored approach achieving 0.04 KL divergence alongside Zopfli compression.

## BACKGROUND

Quantization reduces the bit-precision of model weights (such as converting 32-bit floating-point numbers to smaller integer formats) to lower memory usage and speed up calculation. Quantization-Aware Training (QAT) simulates these low-precision constraints during fine-tuning so the model can learn to preserve accuracy despite loss of numerical precision.

## REFERENCES

## KEYWORDS

#Model Compression#Quantization#Local LLM#JavaScript#Edge AI

$ subscribe --daily

A Developer Compresses a Local Story-Generating LLM into a 24KB HTML File | Daily News