Oído: Open-Source Speech Recognition Outperforming Whisper-Tiny on a $5 Microcontroller
The Lokutor team released Oído, an open-source speech recognition system that runs an 8-bit quantized 13-million parameter Conformer-CTC model directly on an inexpensive ESP32-S3 microcontroller. It achieves better accuracy than OpenAI's Whisper-tiny running on a laptop, scoring lower Word Error Rates (WER) on standard and noisy speech benchmarks. This project demonstrates that high-quality, real-time automatic speech recognition can be deployed on low-cost hardware without relying on dedicated GPUs, NPUs, or cloud APIs. It opens up new possibilities for fully local, privacy-preserving voice interfaces in smart home devices and embedded IoT hardware. Oído uses an 8-bit quantized NVIDIA Conformer-CTC Small model executing on an ESP32-S3 chip equipped with 8 MB PSRAM, requiring no dedicated GPU or NPU accelerator. On the LibriSpeech benchmark, it achieved clean/other WERs of 3.7 / 8.2 compared to Whisper-tiny's 6.3 / 15.9, while also maintaining a lower mean WER (8.4 vs 12.1) under realistic noisy conditions.
## BACKGROUND
Automatic Speech Recognition (ASR) performance is typically evaluated using Word Error Rate (WER), where a lower percentage indicates fewer transcription mistakes. Popular ASR models like OpenAI's Whisper use autoregressive transformer architectures that require substantial computing power and memory. Conformer-CTC architectures combine convolutional layers and self-attention with Connectionist Temporal Classification, offering non-autoregressive decoding that is drastically more efficient for resource-constrained edge microcontrollers.