RESEARCH PAPER

VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs

Haoran Yuan; Weigang Yi; Zhenyu Zhang; Wendi Chen; Yuchen Mo; Jiashi Yin; Xinzhuo Li; Xiangyu Zeng; Chuan Wen; Cewu Lu; Katherine Driggs-Campbell; Ismini Lourentzou

Classification

View four quadrants
Major category
WAMs
Quadrant
Not assigned
Architecture
Not assigned
Prediction paradigm
Not assigned
Source review status
Not assigned

Category review. A visuotactile predictive backbone conditions a separate action expert that generates executable robot controls along with state and force-proxy predictions. The full system is a contact-rich predictive controller. Reading evidence

AT A GLANCE

Contribution

VTAM adapts a video world model to predict camera and tactile streams, then trains a conditional action expert with an auxiliary deformation-derived force target. It reports large gains in three real robot contact tasks. The strongest evidence is task success and a chip ablation; calibrated force accuracy, broad generalization and the claimed gradient mechanism remain unestablished. Method details below preserve inconsistencies between the formulation and implementation appendix.

Abstract

An abstract has not been added yet.

Affiliations

University of Illinois Urbana-Champaign; Stanford University; Shanghai Jiao Tong University