XPeng to Roll Out Second-Generation VLA Model for Autonomous Driving in September
XPeng announced the September rollout of its second-generation Vision-Language-Action (VLA) model for Ultra and Ultra SE vehicles, alongside a distilled "Lite" version for lower-compute platforms. The update introduces the X-Foresight predictive world model and the Infini-VLA long-context architecture. This deployment represents a significant real-world application of advanced multimodal AI in production vehicles, bridging L2 to L4 autonomous driving capabilities. By utilizing model distillation, XPeng demonstrates how high-end AI models can be optimized for edge deployment on vehicles with limited computing power. The new VLA model features token compression to maintain visual understanding with fewer tokens, while the X-Foresight model can predict traffic behaviors 6 seconds into the future. Additionally, the Infini-VLA architecture retains a 30-second memory window of the surrounding environment to improve driving decisions.
## BACKGROUND
A Vision-Language-Action (VLA) model is a multimodal AI that processes visual inputs and instructions to directly output control actions. In autonomous driving, world models are used to predict and simulate future traffic scenarios, helping the vehicle make safer decisions by anticipating the actions of other road users.