XPeng Upgrades Second-Gen VLA Model with 3x Faster Response, Debuting on G9L
XPeng has announced the first major upgrade to its second-generation Vision-Language-Action (VLA) model, version 630, which will debut globally on the XPeng G9L vehicle. The upgrade increases on-device model parameters by 3.5 times and boosts end-to-end response speed by 300%. This upgrade significantly enhances the decision-making and safety capabilities of autonomous driving systems by deploying larger, more capable models directly on the vehicle's edge hardware. The 300% response speed improvement addresses a critical latency bottleneck in end-to-end autonomous driving, enabling smoother and safer real-time maneuvers. XPeng claims that the upgraded on-device model's parameter size is 15 times larger than mainstream VLA models in the industry. The increased response speed is expected to make safety negotiations and deceleration operations smoother during complex driving scenarios.
## BACKGROUND
A Vision-Language-Action (VLA) model is a multimodal foundation model that integrates visual perception, natural language understanding, and physical actions. Originally popularized in robotics, VLA models are increasingly adapted for autonomous driving to translate camera feeds and driving goals directly into vehicle control commands. By processing these models directly on-device (at the edge), vehicles can make faster, more secure decisions without relying on cloud connectivity.