audio.cpp Updates Lower Higgs Audio TTS VRAM by 48% and Accelerate Audio Models
The open-source audio.cpp project released major runtime optimizations, cutting peak VRAM usage for Higgs Audio TTS by 48% down to around 6 GB without accuracy trade-offs. The update also boosts HTDemucs audio separation speed by up to 2.21× on CUDA and accelerates PocketTTS by 2.23× on CPU. By substantially lowering memory requirements and boosting speeds across backends like CUDA, Vulkan, and CPU, audio.cpp makes running state-of-the-art audio models feasible on consumer hardware. This expands the ecosystem for local, privacy-focused AI beyond text generation into expressive text-to-speech and music source separation. The optimizations retain complete model output parity across 110+ supported audio model families and 190+ variants on CUDA, Vulkan, Metal, AMD/HIP, and CPU. Additionally, the update adds an experimental generation history feature to the WebUI that lets users review past outputs and restore their generation settings.
## BACKGROUND
audio.cpp is an open-source framework for running audio models locally with high efficiency, inspired by tools like llama.cpp for text models. It supports diverse audio tasks, including conversational text-to-speech foundation models like Boson AI's Higgs Audio, lightweight offline speech tools like PocketTTS, and music stem separation models like Hybrid Transformer Demucs (HTDemucs).