AGIBOT's WITA-Omni Preview Tops DailyOmni Embodied Multimodal Benchmark
AGIBOT's self-developed embodied native multimodal model, WITA-Omni Preview, has topped the DailyOmni benchmark with a score of 85.21. It outperformed leading models from Qwen, Gemini, Doubao, and NVIDIA, ranking first in six of the eight sub-metrics. The model introduces a "Thinker-Talker-Actor" architecture that elevates physical actions and facial expressions to primary outputs alongside speech. This unified end-to-end approach allows humanoid robots to seamlessly integrate perception, decision-making, and physical execution in real-world environments. WITA-Omni Preview was trained on tens of millions of hours of multimodal data using a hierarchical data system. To address the lack of interaction data in public datasets, AGIBOT constructed a high-quality dataset that preserves the natural temporal relationships between sound, vision, language, actions, and expressions.
## BACKGROUND
The DailyOmni benchmark evaluates a model's ability to perform synchronous cross-modal reasoning and temporal alignment using real-world audio-visual scenarios. Traditional multimodal architectures often use a "Thinker-Talker" paradigm, which decouples reasoning from speech generation, but lack the direct integration of physical actions required for embodied AI.