SV-WAM: An Efficient Surround-View World-Action Model for End-to-End Autonomous Driving
Classification
View four quadrants- Major category
- WAMs
- Architecture
- One Model
- Prediction paradigm
- Joint prediction
- Subcategories
- Autonomous drivingJoint video-action modeling
- Source review status
- Not assigned
Category review. Shared DiT blocks jointly learn action and future-video prediction for surround-view driving. An asymmetric mask allows action-only inference without removing the training-time world-action coupling. Shared DiT blocks jointly learn actions and future video with an asymmetric mask; future rollout is optional at inference. Reading evidence
Contribution
SV-WAM trains one shared transformer to denoise driving actions and future surround-view video, but blocks action attention to future-video tokens. Deployment retains six-camera history while generating only trajectories. A differentiable vehicle-footprint loss improves road compliance. The strongest reported result is 91.0 EPDMS on NAVSIMv2 navtest; the evidence concerns benchmark planning, with weaker hard-split performance and no demonstrated real-vehicle deployment.
Abstract
An abstract has not been added yet.
Affiliations
Institute of Automation, Chinese Academy of Sciences; Chongqing Changan Technology Co., Ltd.; Civil Aviation University of China; Beihang University; Guilin University of Electronic Technology