VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs
Classification
View four quadrants- Major category
- WAMs
- Quadrant
- Not assigned
- Architecture
- Not assigned
- Prediction paradigm
- Not assigned
- Subcategories
- Multimodal, tactile & audio sensing
- Source review status
- Not assigned
Category review. A visuotactile predictive backbone conditions a separate action expert that generates executable robot controls along with state and force-proxy predictions. The full system is a contact-rich predictive controller. Reading evidence
Contribution
VTAM adapts a video world model to predict camera and tactile streams, then trains a conditional action expert with an auxiliary deformation-derived force target. It reports large gains in three real robot contact tasks. The strongest evidence is task success and a chip ablation; calibrated force accuracy, broad generalization and the claimed gradient mechanism remain unestablished. Method details below preserve inconsistencies between the formulation and implementation appendix.
Abstract
An abstract has not been added yet.
Affiliations
University of Illinois Urbana-Champaign; Stanford University; Shanghai Jiao Tong University