WorldLens: Full-Spectrum Evaluations of Driving World Models in Real World
Classification
View four quadrants- Major category
- Benchmarks & simulators
- Quadrant
- Not applicable
- Architecture
- Not applicable
- Prediction paradigm
- Not applicable
- Subcategories
- World model benchmarksAutonomous driving benchmarks
- Source review status
- Not assigned
Category review. This benchmark measures generated driving worlds with reconstruction, perception, human ratings and external planners in closed-loop simulation; its critic outputs judgments rather than driving actions. The cataloged contribution is a dataset/data-generation method or evaluation/simulation resource, not the architecture of an evaluated or external policy. Reading evidence
Contribution
WorldLens evaluates driving video generators through appearance, reconstructability, planner behavior, perception and human judgment. Its strongest lesson is that favorable image metrics coexist with poor closed-loop route completion. A separate LoRA-trained critic learns score-and-rationale outputs from human annotations; its generalization evidence remains qualitative.
Abstract
An abstract has not been added yet.
Affiliations
Not listed in the collection.