Suno Launches 'Speech' Beta to Generate Spoken Audio and Music Together
AI music creation platform Suno has released "Speech," an end-to-end model that generates spoken voice and original background music together as a single audio track. Currently available in public beta, the feature lets users input text alongside prompt descriptions for voice tone and musical style. By replacing traditional multi-step workflows that layer Text-to-Speech (TTS) audio over separate background tracks, Suno simplifies content creation for storytelling, poetry, and podcasts. This launch highlights an ongoing shift in generative AI toward unified models capable of multi-layered audio synthesis. Suno noted several known limitations in the current beta release, including mid-track accent drifting (such as shifting between British and Australian accents) and overly dramatic pauses that slow down pacing. The feature was made public after a month of restricted testing with a small user group.
## BACKGROUND
Conventional audio production typically requires generating a spoken voice track using Text-to-Speech (TTS) software and subsequently mixing it with background music in a digital audio workstation. In contrast, end-to-end generative audio models train a unified neural network to synthesize both spoken dialogue and accompanying orchestration simultaneously from text inputs.