~/AUTONOMOUS D/xpeng-rolls-out-xos-6-3-0-featuring-2nd-gen-vla-and

XPeng Rolls Out XOS 6.3.0 Featuring 2nd-Gen VLA and On-Vehicle World Model

XPeng has officially begun rolling out XOS 6.3.0, introducing its second-generation Vision-Language-Action (VLA) model that transitions autonomous driving from 3D spatial recognition to 4D spatio-temporal understanding. The update includes the X-Foresight predictive world model, which forecasts traffic trajectories up to 6 seconds into the future, as well as a lightweight VLA Lite model for lower-compute hardware. Deploying long-sequence spatio-temporal models and world prediction directly on edge vehicles marks a significant evolution toward higher-level autonomous driving (L2 to L4). XPeng's use of token compression and model distillation also demonstrates how high-capacity AI architectures can be efficiently scaled to mass-production vehicles with hardware constraints. The update introduces the Infini-VLA long-context architecture, which retains 30 seconds of historical temporal context to inform driving decisions. Furthermore, XPeng applied learned token compression and model distillation to deploy the Turing VLA 2.0 Lite model onto single-Turing-chip vehicle configurations without losing visual understanding capacity.

## BACKGROUND

In autonomous driving, Vision-Language-Action (VLA) models combine visual perception, natural language reasoning, and action planning into end-to-end AI pipelines. World models simulate physical reality to predict how environments will evolve, while model distillation transfers knowledge from a large teacher model to a smaller, lightweight model capable of running efficiently on edge compute hardware.

## REFERENCES

## KEYWORDS

#Autonomous Driving#VLA Models#World Models#Edge AI#Computer Vision

$ subscribe --daily

XPeng Rolls Out XOS 6.3.0 Featuring 2nd-Gen VLA and On-Vehicle World Model | Daily News