RESEARCH PAPER
DriveDreamer-Policy: A Geometry-Grounded World-Action Model for Unified Generation and Planning
Classification
View four quadrants- Major category
- WAMs
- Quadrant
- Not assigned
- Architecture
- Not assigned
- Prediction paradigm
- Not assigned
- Subcategories
- Autonomous drivingJoint video-action modeling
- Source review status
- Not assigned
Category review. Shared multimodal query features connect future depth/video supervision to a generative trajectory expert. Planning uses future-aware query context without requiring rendered world outputs. Reading evidence
Contribution
DriveDreamer-Policy trains a shared multimodal backbone with depth, video and trajectory generators. Ordered query embeddings transfer geometry and future-scene context into planning without requiring rendered depth or video at inference. Navsim scores and controlled modality ablations support useful joint supervision, while depth evaluation against a learned teacher and incomplete runtime details limit the conclusions.
Abstract
An abstract has not been added yet.
Affiliations
GigaAI; University of Toronto; CUHK MMLab