Vega: Learning to Drive with Natural Language Instructions
Classification
View four quadrants- Major category
- WAMs
- Quadrant
- Not assigned
- Architecture
- Not assigned
- Prediction paradigm
- Not assigned
- Subcategories
- Autonomous drivingJoint video-action modeling
- Source review status
- Not assigned
Category review. Coupled modality-specific transformers train trajectory generation with action-conditioned future-image prediction; shared global attention integrates the action and world tasks. Reading evidence
Contribution
Vega learns instruction-conditioned driving by training trajectory denoising together with future-image denoising. Modality-specific transformers exchange information through global causal attention. InstructScene supplies automatically generated instructions describing recorded driving. The strongest reported NAVSIM v2 score uses best-of-six trajectory selection; NAVSIM v1 results are weaker than leading VLA baselines. Future-prediction ablations support visual supervision, while selected images illustrate instruction sensitivity without measuring general instruction-following reliability.
Abstract
An abstract has not been added yet.
Affiliations
Tsinghua University; GigaAI