~/AUDIO LANGUA/fireredteam-releases-fireredaudio-and-fireredtts3-for-unified-audio-understanding-and-generation

FireRedTeam Releases FireRedAudio and FireRedTTS3 for Unified Audio Understanding and Generation

FireRedTeam has released FireRedAudio, a 9B-parameter general-purpose audio language model, alongside FireRedTTS3, a unified speech generation and editing system. FireRedAudio features a novel architecture that decouples understanding and generation pathways while sharing a single LLM backbone. This release provides the open-source community with a unified model capable of handling the entire audio stack, including speech recognition, zero-shot text-to-speech, and speech editing. Its decoupled representation design prevents interference between understanding and generation tasks, setting a new standard for audio-language models. FireRedAudio supports precise temporal grounding for audio files up to one hour long. FireRedTTS3 supports zero-shot voice cloning across 24 languages (including Vietnamese) and 21 Chinese dialects, as well as instruction-driven voice design.

## BACKGROUND

Traditional audio models often separate speech-to-text (understanding) and text-to-speech (generation) into distinct systems, or struggle to balance both in a single model. Temporal grounding is the process of aligning specific timestamps in an audio recording with its corresponding textual or semantic content. Decoupled continuous representations help resolve this by using separate pathways for encoding and decoding while leveraging a shared reasoning model.

## REFERENCES

## KEYWORDS

#Audio Language Models#Text-to-Speech#Speech Recognition#Open Source AI

$ subscribe --daily

FireRedTeam Releases FireRedAudio and FireRedTTS3 for Unified Audio Understanding and Generation | Daily News