~/AUTONOMOUS D/xpeng-announces-second-generation-vla-model-with-4d-spacetime-understanding

XPeng Announces Second-Generation VLA Model with 4D Spacetime Understanding

XPeng has unveiled its second-generation Vision-Language-Action (VLA) model, which transitions from 3D spatial perception to 4D spacetime understanding by integrating the time dimension. The updated model features the "Infini-VLA" long-sequence architecture, allowing it to remember the previous 30 seconds of driving context. By incorporating temporal context and boosting end-to-end response speeds by 300%, this upgrade enables autonomous vehicles to make faster, safer, and more continuous decisions in dynamic real-world environments. This marks a significant step forward in physical AI, where models must perceive and act in real-time rather than processing static frames. The second-generation VLA utilizes streaming autoregressive inference to achieve parallel processing of perception, reasoning, and action. However, the larger model size and longer temporal context demand significantly higher computational power and faster processing to maintain real-time safety.

## BACKGROUND

A Vision-Language-Action (VLA) model is a multimodal AI that integrates visual inputs and language instructions to directly generate physical actions for robots or autonomous systems. Physical AI refers to systems that interact with the physical world through sensors and actuators, distinguishing them from purely digital generative AI.

## REFERENCES

## KEYWORDS

#Autonomous Driving#VLA Model#Embodied AI#Physical AI

$ subscribe --daily

XPeng Announces Second-Generation VLA Model with 4D Spacetime Understanding | Daily News