RESEARCH PAPER

Vega: Learning to Drive with Natural Language Instructions

Sicheng Zuo; Yuxuan Li; Wenzhao Zheng; Zheng Zhu; Jie Zhou; Jiwen Lu

Classification

View four quadrants
Major category
WAMs
Quadrant
Not assigned
Architecture
Not assigned
Prediction paradigm
Not assigned
Source review status
Not assigned

Category review. Coupled modality-specific transformers train trajectory generation with action-conditioned future-image prediction; shared global attention integrates the action and world tasks. Reading evidence

AT A GLANCE

Contribution

Vega learns instruction-conditioned driving by training trajectory denoising together with future-image denoising. Modality-specific transformers exchange information through global causal attention. InstructScene supplies automatically generated instructions describing recorded driving. The strongest reported NAVSIM v2 score uses best-of-six trajectory selection; NAVSIM v1 results are weaker than leading VLA baselines. Future-prediction ablations support visual supervision, while selected images illustrate instruction sensitivity without measuring general instruction-following reliability.

Abstract

An abstract has not been added yet.

Affiliations

Tsinghua University; GigaAI