Wow, wo, val! A Comprehensive Embodied World Model Evaluation Turing Test
Classification
View four quadrants- Major category
- Benchmarks & simulators
- Quadrant
- Not applicable
- Architecture
- Not applicable
- Prediction paradigm
- Not applicable
- Source review status
- Not assigned
Category review. The work contributes curated embodied evaluation cases, automated/human video judgments and a separate IDM-based execution test for existing generators. The cataloged contribution is a dataset/data-generation method or evaluation/simulation resource, not the architecture of an evaluated or external policy. Reading evidence
Contribution
WoW-World-Eval tests whether instruction-conditioned robot videos are visually convincing, task-correct, physically plausible and usable for action extraction. Its 609-sample benchmark combines automated metrics, human judgments and a separate GC-IDM robot-execution test. Hailuo leads the reported aggregate video score, while WoW-wan leads physical execution. The central lesson is that a plausible imagined manipulation and an executable one are distinct outcomes; calibration and evaluator dependence constrain how broadly the scores can be interpreted.
Abstract
An abstract has not been added yet.
Affiliations
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University; Beijing Innovation Center of Humanoid Robotics; The Hong Kong University of Science and Technology