RESEARCH PAPER

Compact Visuotactile World Models for Lifting: Prediction, Reward Alignment, and Force Constraints

Qinzhen Ma; Sida Peng

Classification

View four quadrants
Major category
WAMs
Architecture
Pending verification
Prediction paradigm
Other mechanisms
Source review status
Not assigned

Category review. An action-conditioned visuotactile latent model supplies imagined returns for learning executable lifting actions; this is task-specific model-assisted control, not a reusable generic component. The described action mechanism is external planning, model-assisted policy optimization, geometric tracking, or video-conditioned control; the source does not establish joint future/action generation or an IDM action decoder. Reading evidence

AT A GLANCE

Contribution

This study asks whether better contact prediction produces better lifting decisions. A compact visuotactile model improves forecasts over vision alone, but persistence challenges the value of its dynamics. An exploratory reward revision improves executed simulator lifting, while simple force feedback remains stronger under a force budget. A separate GelSight experiment exposes the gap between frame and trajectory reliability.

Abstract

An abstract has not been added yet.

Affiliations

Rice University; Zhejiang University