Compact Visuotactile World Models for Lifting: Prediction, Reward Alignment, and Force Constraints
Classification
View four quadrants- Major category
- WAMs
- Quadrant
- Outside quadrants
- Architecture
- Pending verification
- Prediction paradigm
- Other mechanisms
- Source review status
- Not assigned
Category review. An action-conditioned visuotactile latent model supplies imagined returns for learning executable lifting actions; this is task-specific model-assisted control, not a reusable generic component. The described action mechanism is external planning, model-assisted policy optimization, geometric tracking, or video-conditioned control; the source does not establish joint future/action generation or an IDM action decoder. Reading evidence
Contribution
This study asks whether better contact prediction produces better lifting decisions. A compact visuotactile model improves forecasts over vision alone, but persistence challenges the value of its dynamics. An exploratory reward revision improves executed simulator lifting, while simple force feedback remains stronger under a force budget. A separate GelSight experiment exposes the gap between frame and trajectory reliability.
Abstract
An abstract has not been added yet.
Affiliations
Rice University; Zhejiang University