~/ROBOTICS/unitree-releases-unifolm-wla-1-0-a-6b-model-for-humanoid-whole

Unitree Releases UnifoLM-WLA-1.0, a 6B Model for Humanoid Whole-Body Control

Unitree Robotics has released UnifoLM-WLA-1.0, a 6-billion parameter general-purpose humanoid foundation model trained on approximately 2,500 hours of real-robot data. The single model unifies whole-body control and tabletop manipulation to accomplish 64 distinct physical tasks on humanoid platforms such as the Unitree G1. This release marks a significant step forward in embodied AI by demonstrating that a single model can handle both high-level spatial reasoning and fine-grained whole-body control across varied end-effectors. It advances open-source humanoid robotics toward practical generalist automation rather than fragmented, single-task control systems. Built upon the UnifoLM-ER-1 embodied reasoner (derived from Qwen3-VL), the model combines optical flow and VQ-VAE for dynamic region prediction alongside Residual Vector Quantization for action discretization. A Multimodal Diffusion Transformer (MMDiT) action expert layer is then integrated to generate continuous control signals for lower-body movement, arms, and dexterous hands.

## BACKGROUND

Vision-Language-Action (VLA) models are multimodal foundational networks that interpret visual observations and text instructions to directly output physical robotic actions. Historically, humanoid robotics relied on separate control stacks for balancing, navigation, and object manipulation, making general-purpose task execution difficult to unify into a single model.

## REFERENCES

## KEYWORDS

#Robotics#Embodied AI#VLA Models#Computer Vision#Multimodal AI

$ subscribe --daily

Unitree Releases UnifoLM-WLA-1.0, a 6B Model for Humanoid Whole-Body Control | Daily News