RESEARCH PAPER

Dream-Tac: A Unified Tactile World Action Model for Contact-Rich Robot Manipulation

Lou, Yunfan; Ye, Yifan; Fu, Yankai; Cen, Jun; Chi, Xiaowei; Lyu, Yaoxu; Jia, Peidong; Han, Sirui; Lu, Zhihe; Zhang, Shanghang

Classification

View four quadrants
Major category
WAMs
Architecture
One Model
Prediction paradigm
Joint prediction
Source review status
Verified from primary sources
AT A GLANCE

Contribution

In this paper, we propose Dream-Tac, a unified Tactile-World Action Model that jointly models actions, future visual observations, and tactile dynamics.

Abstract

World action models inherit the predictive capability of world models, enabling action generation to be guided by anticipated future observations. However, they rely primarily on vision and often fail in contact-rich manipulation, where critical cues arise from physical interaction. In this paper, we propose Dream-Tac, a unified Tactile-World Action Model that jointly models actions, future visual observations, and tactile dynamics. Specifically, Dream-Tac introduces (i) contact-gated visuotactile fusion to selectively integrate tactile signals and (ii) a contact-aware attention bias to better regulate cross-modal interactions during manipulation. To support real-time deployment, we further design a dual-level acceleration strategy, reformulating the contact-aware bias to preserve the fused attention path during training and introducing cache-based diffusion acceleration at inference, achieving up to 2.9×\times faster training and 1.8×\times faster inference. Across six contact-rich manipulation tasks, Dream-Tac improves action accuracy by 31.7\% on average, demonstrating the effectiveness of unified visuotactile world modeling.Code is available at https://github.com/LYFCLOUDFAN/Dream-Tac.

Affiliations

Peking University; Beijing; China; The Hong Kong University of Science and Technology; Hong Kong; Nanjing University; Nanjing; State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University