ManipArena: A Controlled Benchmark for Diagnosing Generalization in Real-Robot Manipulation
Classification
View four quadrants- Major category
- Benchmarks & simulators
- Quadrant
- Not applicable
- Architecture
- Not applicable
- Prediction paradigm
- Not applicable
- Source review status
- Not assigned
Category review. The primary contribution is controlled real-robot evaluation with task schemas, subgoal rubrics and generalization splits for existing VLA/WAM policies. The cataloged contribution is a dataset/data-generation method or evaluation/simulation resource, not the architecture of an evaluated or external policy. Reading evidence
Contribution
ManipArena evaluates manipulation policies through controlled physical robot trials, task schemas and subgoal scoring. Its strongest lesson is that training recipes and model provenance affect rankings alongside architecture. Language grounding and demonstration selection produce substantial reported gains, but small trial counts, restricted environments and internal reporting inconsistencies limit causal and generalization claims.
Abstract
An abstract has not been added yet.
Affiliations
Sun Yat-sen University; X Square Robot; MBZUAI; Tsinghua University; University of Zurich