~/SPEECH TO TE/cactus-compute-releases-whistle-a-16-9mb-ultra-compact-speech-to-text

Cactus Compute Releases Whistle: A 16.9MB Ultra-Compact Speech-to-Text Model

Cactus Compute has introduced Cactus Whistle, an open 55-million parameter (36-million active) speech-to-text model compressed to just 16.9MB using CQ2bit quantization. Despite being 9x smaller and 6x faster than OpenAI's Whisper base, Whistle outperforms it across standard benchmarks including LibriSpeech and FLEURS. This development enables high-accuracy multi-lingual automatic speech recognition directly on edge devices such as microcontrollers, wearables, and budget smartphones without needing cloud processing. It demonstrates that targeted architectural innovations and aggressive model quantization can make powerful AI capabilities accessible on resource-constrained hardware. Whistle combines a log-mel audio encoder stem with a Simple Attention and Hadamard MLP decoder featuring gated cross-attention. It natively supports keyword biasing during beam search to accurately transcribe uncommonly spelled names, provides word-level timestamps, and is deployable across 17 platforms including RISC-V, ARM64, and WebAssembly.

## BACKGROUND

Automatic Speech Recognition (ASR) models like OpenAI's Whisper convert spoken audio into written text, but standard models often require hundreds of megabytes of memory and heavy computing resources. Model quantization reduces memory footprint by converting model weights to lower-precision bit representations, allowing complex neural networks to run locally on low-power hardware.

## REFERENCES

## KEYWORDS

#Speech-to-Text#ASR#Edge AI#Model Quantization#Machine Learning

$ subscribe --daily

Cactus Compute Releases Whistle: A 16.9MB Ultra-Compact Speech-to-Text Model | Daily News