~/AUTONOMOUS D/xpeng-unveils-second-generation-vla-model-and-master-agent-for-autonomous-vehicles

XPeng Unveils Second-Generation VLA Model and Master Agent for Autonomous Vehicles

XPeng has upgraded its second-generation Vision-Language-Action (VLA) model and introduced the "Master Agent," which integrates VLA and Vision-Language Models (VLM) to reconstruct the vehicle's brain. This system allows the vehicle to understand complex user intents and execute physical actions, such as pulling over via voice commands. This deployment represents a major milestone in Embodied AI, bridging the gap between large language model reasoning and real-world physical vehicle control. By orchestrating driving, cabin, and chassis systems, it moves autonomous driving closer to intuitive, human-like interaction and execution. Built on XPeng's self-developed Omni multi-modal model, the Master Agent can decompose vague user requests into tasks and coordinate vertical agents across driving, chassis, cabin, and body control. For example, it can locate a destination based on a vague description of a storefront and plan the route accordingly.

## BACKGROUND

A Vision-Language-Action (VLA) model is a multimodal foundation model that integrates visual perception, natural language understanding, and physical action generation. Embodied AI refers to artificial intelligence that interacts directly with the physical world through sensors and actuators, transitioning AI from purely digital reasoning to physical execution.

## REFERENCES

## KEYWORDS

#Autonomous Driving#VLA Models#AI Agents#Embodied AI#Smart Cabin

$ subscribe --daily

XPeng Unveils Second-Generation VLA Model and Master Agent for Autonomous Vehicles | Daily News