~/MULTIMODAL A/bytedance-launches-seedrealtime-a-native-audio-video-full-duplex-ai-model

ByteDance Launches SeedRealtime, a Native Audio-Video Full-Duplex AI Model

ByteDance's Seed team has released SeedRealtime, a native audio-video full-duplex large model that integrates audio, video, and text into a unified architecture. The model enables real-time, multimodal interaction and has been fully deployed on the Doubao app. This launch represents a shift from traditional turn-based AI interactions to natural, simultaneous "look, listen, and speak" communication. By scaling this technology directly to the Doubao app, ByteDance marks a major milestone in practical, real-time human-computer interaction. SeedRealtime reduces conversational rhythm issues by half compared to traditional cascade models, significantly improving the flow of dialogue and reducing interruptions. Users can access this feature by updating the Doubao app and selecting the video call option.

## BACKGROUND

Traditional voice assistants typically use a "cascade" architecture that chains separate speech-to-text, language processing, and text-to-speech models, resulting in lag and turn-based interactions. In contrast, full-duplex models allow for continuous, simultaneous two-way communication, similar to a natural human conversation. The ByteDance Seed team, established in 2023, focuses on general intelligence research including multimodal AI and next-generation interactions.

## REFERENCES

## KEYWORDS

#Multimodal AI#Large Language Models#Human-Computer Interaction#ByteDance

$ subscribe --daily

ByteDance Launches SeedRealtime, a Native Audio-Video Full-Duplex AI Model | Daily News