RESEARCH PAPER

4DGS-WAM: Bridging Past and Future with an Object-Centric World Action Model based on 4D Gaussian Splatting

Ma, Yueen; Xu, Zenglin; King, Irwin

Classification

View four quadrants
Major category
WAMs
Architecture
Dual-system
Prediction paradigm
Other mechanisms
Source review status
Verified from primary sources

Category review. A concrete 2026 object-centric system couples a learned motion policy to an action-conditioned Gaussian world model. This is a specialized world-action method, not a historical foundation or a generic geometry component. Separate policy and dynamics networks are trained separately; actor motions precede future rendering. Actions are actor/ego motion forecasts, not demonstrated executed robot commands. KITTI image/reconstruction evaluation does not establish closed-loop control. Reading evidence

AT A GLANCE

Contribution

These models can achieve exceptional visual quality, but they lack explicit spatial structure for individual objects and repeatedly process redundant background content.

Abstract

Current world action models (WAMs) typically operate on 2D visual data. These models can achieve exceptional visual quality, but they lack explicit spatial structure for individual objects and repeatedly process redundant background content. Although point clouds can represent the world in 3D space, they can be difficult to align and accumulate across viewpoints. In this paper, we leverage an explicit 4D Gaussian Splatting (4DGS) representation that separately models dynamic objects and the static background of a scene. For dynamic objects, we use a policy model to predict future actor actions and a world model to predict transformations of their observed Gaussian splats. The static background need not be regenerated for future states, as much of it has already been observed in past frames. This forms an object-centric world action model, which we name 4DGS-WAM. It lifts 2D observations into a persistent 4D representation so that previously observed static content can be reused during future prediction. Future-state extrapolation can then focus on modeling the evolution of dynamic objects. Experiments on KITTI-MOT evaluate short-horizon prediction and past reconstruction.

Affiliations

The Chinese University of Hong Kong; Shanghai Academy of AI for Science; Fudan University

BibTeX