XPeng Enters Driverless Testing Phase with Second-Gen VLA Model Unifying L2-L4
XPeng Motors has entered the driverless road testing phase in Guangzhou for its Robotaxi and announced that its second-generation Vision-Language-Action (VLA) model has unified L2 to L4 autonomous driving capabilities. The model can also be distilled into a lighter version (Turing VLA 2.0 Lite) to run on lower-compute hardware platforms. This development demonstrates how physical AI foundation models can bridge consumer-grade driver assistance (L2) and fully autonomous Robotaxis (L4). By deploying distilled versions of a unified model, XPeng can scale advanced autonomous driving capabilities across a wider range of mass-production vehicles. The second-generation VLA achieves high-quality visual understanding with fewer tokens through learned token compression and distillation training, preserving model capacity while reducing computational overhead. This allows the model to generalize globally and adapt to lower-compute platforms.
## BACKGROUND
A Vision-Language-Action (VLA) model is a multimodal foundation model that translates visual inputs and instructions directly into physical actions. Physical AI refers to AI systems, like autonomous vehicles and robots, that interact directly with the physical world. Autonomous driving levels range from L2 (partial automation requiring driver supervision) to L4 (high automation where the vehicle can drive itself under specific conditions).