RESEARCH PAPER

RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation

Feng Jiang; Yang Chen; Kyle Xu; Yuchen Liu; Haifeng Wang; Zhenhao Shen; Jasper Lu; Shengze Huang; Yuanfei Wang; Chen Xie; Ruihai Wu

Classification

View four quadrants
Architecture
Not applicable
Prediction paradigm
Not applicable
Source review status
Not assigned

Category review. The contribution is an execution-based benchmark that converts generated robot or human videos to actions through separate interfaces and scores task completion in a simulator. The cataloged contribution is a dataset/data-generation method or evaluation/simulation resource, not the architecture of an evaluated or external policy. Reading evidence

AT A GLANCE

Contribution

RoboWM-Bench evaluates whether generated manipulation videos can be converted into robot actions that complete tasks in simulation. Separate human-hand retargeting and robot inverse-dynamics interfaces expose failures hidden by plausible imagery. Reliability tests support these interfaces on real demonstrations, but execution scores still depend on action extraction, reconstructed physics and task-specific checkers.

Abstract

An abstract has not been added yet.

Affiliations

Peking University; Tsinghua University; Lightwheel