~/MULTIMODAL A/inclusionai-releases-realtime-venus-open-source-9b-full-duplex-multimodal-ai-model

InclusionAI Releases Realtime-Venus: Open-Source 9B Full-Duplex Multimodal AI Model

InclusionAI has released Realtime-Venus, an open-source 9B multimodal AI system adapted from MiniCPM-o 4.5 that supports continuous audio-visual perception, streaming text and speech generation, and proactive interaction. The project provides two model checkpoints: Realtime-Venus-Omni for combined audio-visual capabilities and Realtime-Venus-Audio for streaming audio processing. This release advances open-source multimodal AI by providing true full-duplex conversational models capable of listening, seeing, and speaking simultaneously while responding proactively without waiting for user prompts. It moves the open ecosystem closer to natural human-like assistants that handle interruptions, backchannels, and asynchronous task delegation in real time. Realtime-Venus incorporates semantic interruption handling, training-free long-video memory retrieval, and an in-stream `<delegate>` request syntax that passes tasks to background execution runtimes without pausing ongoing speech. Audio synthesis uses bundled Token2wav resources with reference voice support, orchestrated via the Realtime-Venus-Harness repository.

## BACKGROUND

Full-duplex conversational AI allows simultaneous two-way interaction, enabling models to continuously process visual and acoustic inputs while speaking, rather than waiting for discrete user turns. Realtime-Venus builds upon OpenBMB's MiniCPM-o architecture, an end-to-end 9B multimodal model incorporating vision (SigLip), speech recognition (Whisper), and speech generation components.

## REFERENCES

## KEYWORDS

#Multimodal AI#Audio-Visual Models#Open Source#Speech Synthesis#LLM

$ subscribe --daily

InclusionAI Releases Realtime-Venus: Open-Source 9B Full-Duplex Multimodal AI Model | Daily News