RESEARCH PAPER

SV-WAM: An Efficient Surround-View World-Action Model for End-to-End Autonomous Driving

Jinyang Wang; Shiwei Li; Junjian Wang; Zhiqiang Deng; Jianbin Gao; Yihang Zhao; Liu Liu; Yongjia Zhao; Jinlong Chen; Huirui Xu; Yifeng Pan; Kangwei Liu; Fan Ren; Ji Tao; Minghao Yang

Classification

View four quadrants
Major category
WAMs
Architecture
One Model
Prediction paradigm
Joint prediction
Source review status
Not assigned

Category review. Shared DiT blocks jointly learn action and future-video prediction for surround-view driving. An asymmetric mask allows action-only inference without removing the training-time world-action coupling. Shared DiT blocks jointly learn actions and future video with an asymmetric mask; future rollout is optional at inference. Reading evidence

AT A GLANCE

Contribution

SV-WAM trains one shared transformer to denoise driving actions and future surround-view video, but blocks action attention to future-video tokens. Deployment retains six-camera history while generating only trajectories. A differentiable vehicle-footprint loss improves road compliance. The strongest reported result is 91.0 EPDMS on NAVSIMv2 navtest; the evidence concerns benchmark planning, with weaker hard-split performance and no demonstrated real-vehicle deployment.

Abstract

An abstract has not been added yet.

Affiliations

Institute of Automation, Chinese Academy of Sciences; Chongqing Changan Technology Co., Ltd.; Civil Aviation University of China; Beihang University; Guilin University of Electronic Technology