RESEARCH PAPER

MindDrive: An All-in-One Framework Bridging World Models and Vision-Language Model for End-to-End Autonomous Driving

Bin Sun; Yaoguang Cao; Yan Wang; Rui Wang; Jiachen Shang; Xiejie Feng; Jiayi Lu; Jia Shi; Shichun Yang; Xiaoyu Yan; Ziying Song

Classification

View four quadrants
Major category
WAMs
Architecture
Pending verification
Prediction paradigm
Other mechanisms
Subcategories
Autonomous driving
Source review status
Not assigned

Category review. An action-conditioned BEV world model predicts future context for a separate anchor decoder and VLM scorer that refine and select driving trajectories. The described action mechanism is external planning, model-assisted policy optimization, geometric tracking, or video-conditioned control; the source does not establish joint future/action generation or an IDM action decoder. Reading evidence

AT A GLANCE

Contribution

MindDrive couples action-conditioned BEV scene prediction with anchor refinement and a VLM trajectory scorer. Its main contribution is the connection between future-aware candidate generation and selection. NAVSIM scores support planning gains, while incomplete training specifications and inconsistent result reporting limit reproducibility.

Abstract

An abstract has not been added yet.

Affiliations

School of Transportation Science and Engineering, Beihang University; State Key Laboratory of Intelligent Transportation System, Beihang University; Hangzhou International Innovation Institute, Beihang University; Contemporary Amperex Technology Co., Limited (CATL); Research Institute of Aero-Engine, Beihang University; School of Computer Science and Technology, Beijing Jiaotong University; China Automotive Engineering Research Institute Co., Ltd.