RESEARCH PAPER

GaussianWAM: Distilling Geometry and Semantics from 3D Gaussian Fields into World-Action Models

Zhang, Zijian; Jiang, Yuqing; Zhou, Weitao; Li, Minglei; Zhang, Jinhao; Mu, Yao; Li, Xiaofan; Zhao, Hao; Yu, Haibao

Classification

View four quadrants
Major category
WAMs
Architecture
Not applicable
Prediction paradigm
Other mechanisms
Source review status
Verified from primary sources
AT A GLANCE

Contribution

We propose \textbf{GaussianWAM}, a training-time representation-enhancement framework that organizes geometric and semantic supervision through a 3D Gaussian field.

Abstract

World-Action Models (WAMs) jointly learn future visual prediction and action generation, using video dynamics as a representation-learning signal for robotic manipulation. However, their video latents are primarily optimized for visual prediction and are not explicitly encouraged to preserve cross-view geometric structure or spatially localized, object-relevant semantics. We propose \textbf{GaussianWAM}, a training-time representation-enhancement framework that organizes geometric and semantic supervision through a 3D Gaussian field. Given synchronized multi-view observations, frozen geometry and vision foundation models provide depth, camera parameters, and dense semantic features. GaussianWAM binds these heterogeneous signals to shared Gaussian primitives and renders spatially aligned semantic, depth, and coverage targets, which are distilled into the current-observation representations of the WAM. All teacher models, Gaussian components, and auxiliary prediction heads are removed after training, leaving the original WAM inference path without additional modules or forward computation. On LIBERO-Plus, GaussianWAM improves FastWAM from 52.05\% to 71.29\% and Cosmos Policy from 71.52\% to 77.30\%. Direct CLIP and VGGT distillation already establishes a strong FastWAM baseline of 69.37\%, while Gaussian-field unification further improves it to 71.29\%, supporting the benefit of spatially organizing heterogeneous teacher signals. GaussianWAM also improves performance on standard LIBERO and shows positive transfer trends on RoboTwin and real-world manipulation. These results suggest that training-time Gaussian distillation provides a practical way to inject geometry- and semantics-related supervision into WAM representations without changing their deployment architecture.

Affiliations

1Tuojing Intelligence, 2University of Chinese Academy of Sciences,3Institute of Automation, Chinese; Academy of Sciences,4Tsinghua University,5Simple AI, 6Harbin Institute of Technology (Shenzhen); 7Shanghai Jiao Tong University,8Zhejiang University, 9Institute for AI Industry Research (AIR), Tsinghua; University, 10The University of Hong Kong