RESEARCH PAPER

Towards Surgical World-Action Modeling: A Preliminary Joint Visual-Trajectory Forecasting for Surgical Motion Planning

Huang, Weiliang; Liu, Huanrong; Zhang, Bob; Dou, Qi; Chen, Zhen; Gu, Yun; Rosman, Guy; Li, Qingbiao

Classification

View four quadrants
Major category
WAMs
Architecture
One Model
Prediction paradigm
Joint prediction
Source review status
Verified from primary sources

Category review. A 2026 shared temporal-spatial predictor generates future visual latents and surgical-tool trajectories through two heads. This is a specialized joint visual/action-oriented trajectory forecasting method, not a historical foundation. Reading evidence

AT A GLANCE

Contribution

To bridge this gap, we present a preliminary joint visual-trajectory world-action model that simultaneously forecasts future visual states and instrument trajectories from historical surgical observations. The chunked strategy consistently outperforms direct one-shot prediction across all evaluated horizons, improving first-segment PSNR from 18.86 to 23.11 dB and reducing ADE from 45.77 to 22.22 pixels.

Abstract

Reliable surgical planning requires models to anticipate not only how instruments will move, but also how the operative visual state will evolve together with such motion. Existing approaches typically treat future scene generation and instrument trajectory prediction as two separate tasks. Scene-only models cannot directly evaluate the accuracy of future instrument motion at the trajectory level, while trajectory-only models fail to capture the visual consequences of instrument movement, leaving the consistency between predicted trajectories and future scene evolution unaddressed. Jointly forecasting both provides a more complete account of surgical action-scene dynamics by enabling explicit trajectory-level evaluation while simultaneously modeling the corresponding visual evolution. To bridge this gap, we present a preliminary joint visual-trajectory world-action model that simultaneously forecasts future visual states and instrument trajectories from historical surgical observations. Specifically, we encode historical video frames and tool trajectories into latent representations, which are processed by a temporal-spatial encoder and subsequently decoded through separate visual-state and trajectory prediction heads. Based on this preliminary architecture, a chunked autoregressive rollout is repeatedly applied to predict fifteen future steps. The chunked strategy consistently outperforms direct one-shot prediction across all evaluated horizons, improving first-segment PSNR from 18.86 to 23.11 dB and reducing ADE from 45.77 to 22.22 pixels. These results demonstrate the initial feasibility of joint visual-motion forecasting. However, we observe progressive visual degradation and accumulated trajectory errors over longer prediction horizons, which remain important challenges for future surgical world-action modeling.

Affiliations

1 Faculty of Information Science and Computing, University of Macau, Macau, China; 2 Faculty of Engineering, University of Macau, Macau, China; 3 University of Macau Advanced Research Institute in Hengqin, Zhuhai, China; 4 Department of Computer Science and Engineering, The Chinese University of Hong; 5 Department of Data Science and Artificial Intelligence, The Hong Kong Polytechnic; University, Hong Kong, China; 6 Shanghai Key Laboratory of Flexible Medical Robotics, Tongren Hospital, Institute; of Medical Robotics, Shanghai Jiao Tong University, Shanghai, China; 7 School of Automation and Intelligent Sensing, Shanghai Jiao Tong University; 8 School of Medicine, Duke University, Durham, North Carolina, USA