EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models
Classification
View four quadrants- Major category
- Benchmarks & simulators
- Quadrant
- Not applicable
- Architecture
- Not applicable
- Prediction paradigm
- Not applicable
- Subcategories
- World model benchmarks
- Source review status
- Not assigned
Category review. The work curates embodied video-generation cases and evaluates scene, motion and semantic quality of external generators; its detector/encoder/MLLM pipeline returns scores rather than controls. The cataloged contribution is a dataset/data-generation method or evaluation/simulation resource, not the architecture of an evaluated or external policy. Reading evidence
Contribution
EWMBench evaluates instruction-conditioned robot videos through scene consistency, end-effector motion, and semantics. Its seven-model comparison favors domain-adapted generators, while controlled trajectory corruptions expose why static-looking plausibility is insufficient. These are offline video-evaluation results, with unresolved reporting inconsistencies, rather than demonstrations of executed robot control.
Abstract
An abstract has not been added yet.
Affiliations
AgiBot; SJTU; MMLab-CUHK; HIT