Paper library

Papers, datasets, benchmarks, and reading notes on World-Action Models.

564 entries8 major categories1989 – 2026 recorded publication yearsUpdated

564 papers in the collection

Programmable World Model

Zheng-Hui Huang; Guixu Lin; Jiacheng Lin +8 authors

Programmable World Model makes a rule-executing engine authoritative for world facts and uses a conditioned video model to render them. State-augmented 3D oriented bounding boxes connect these components without requiring detailed animated assets. On CombatStateBench, the system reports 94% visible-alive-count accuracy and 98% death-state accuracy, but these permissive global metrics do not establish correct entity-specific interactions or autonomous action selection (e-world, e-controls, e-table, e-metrics).

DUET-DINO: Simultaneous Cross-View World Modeling for Latent Planning in Robot Manipulation

Nisarga Nilavadi; Ralf Römer; Moritz Reuss +5 authors

DUET-DINO couples side- and wrist-camera latent predictors through cross-attention, then searches seven-dimensional robot actions against paired goal images. Its clearest evidence is improved simulated spatial and orientation planning; hardware success remains limited under a smaller planning budget and safety terminations (e-architecture, e-planning, e-angled, e-hardware).

FolDeX: A Physical-World Benchmark for Long-Horizon Robotic Manipulation of Deformable Objects

Chenhuan Liu; Yi Xu; Feng Wu +7 authors

FolDeX evaluates complete physical garment-folding episodes and asks whether expensive robot experience can be reused across recovery states, tasks, scenes, and embodiments. Its reference policies use π0. Recovery-augmented training raises reported average success from 80.75% to 95.00%, but transfer studies remain preliminary and the missing quality-scoring appendix prevents independent reconstruction of FoldScore from the supplied paper alone.

Grounding Generated Video Plans in Simulation Towards Versatile Dexterous Controllers

Tianyue Wu; Boyuan An; Shuqi Zhao +6 authors

GALATEA turns image-and-language-conditioned manipulation videos into 3-D hand–object references, then learns simulation-based controllers that physically track them. Its central interface retains finger motion and object motion together. Multi-skill experts are distilled into one feedback policy; simulated transfer and 27/40 real-world successes on unseen plans support useful, but limited, generalization (e02, e09, e14, e17).

HaWMPO: Hallucination-Aware World Model-based Policy Optimization for Generalist Robot Policy

Zengjue Chen; Peidong Liu; Jiawei Li +1 authors

HaWMPO post-trains OpenVLA-OFT inside a frozen video world model. A separate hallucination detector discounts rewards assigned to unreliable imagined action chunks before GRPO updates the policy. Table 1 reports 63.7% average LIBERO success, but gains vary by suite and the penalty-selection description is inconsistent. Physical testing provides preliminary evidence on two G1 tasks.

JEPA Policy: Diffusion-Free Imitation Learning via Paired Action and Future Representation Prediction

Jie Xu; Kangjin Yu; Ziyi Jin +5 authors

JEPA Policy trains a shared Transformer to predict an action chunk and the visual representation observed later in the same demonstration. Two deterministic passes refine both outputs. Its strongest evidence is improved simulated control relative to action-only MIP, supported by topology controls; future error also offers a delayed, task-dependent rollout diagnostic.

Compact Visuotactile World Models for Lifting: Prediction, Reward Alignment, and Force Constraints

Qinzhen Ma; Sida Peng

This study asks whether better contact prediction produces better lifting decisions. A compact visuotactile model improves forecasts over vision alone, but persistence challenges the value of its dynamics. An exploratory reward revision improves executed simulator lifting, while simple force feedback remains stronger under a force budget. A separate GelSight experiment exposes the gap between frame and trajectory reliability.

Valerant: An Automatic Navigable Game Map Generator via Action-Conditioned World Model Exploration

Yiran Qiao; Feng Wang; Jing Ma

Valerant expands one game screenshot into a persistent 3D point-cloud map by generating alternative action-conditioned videos, reconstructing each with SLAM, and committing the best admissible branch. A frozen video model supplies futures; an external exploration policy selects actions. Perceptual comparison and qualitative ablations support this prototype, while geometric reliability and reproducibility remain incompletely measured.

SyncWorld: Visual Calibration Enables World Models as Zero-Shot Simulators

Yuncong Yang; Zhengtao Han; Furkan Ozyurt +6 authors

SyncWorld interprets robot commands through a short visual calibration context, then predicts proposed actions’ consequences with a video diffusion model. Calibration-conditioned training and distillation also support history-only prediction. It improves held-out video metrics and selected LIBERO policy outcomes, but the evidence is strongest for short-horizon simulation: ranking quality, object interactions and unseen-embodiment controllability remain limiting factors.

WorldAgen: Unified State-Action Prediction with Test-Time World Model Training

Chi Wan; Kangrui Wang; Yuan Si +2 authors

WorldAgen shares a Transformer between action prediction and future-observation prediction, then adapts the shared representation using exploratory target-environment transitions. Simulated manipulation improves after observation-only test-time training (TTT), while implementation ambiguities and unreported uncertainty limit the strength of the generalization claim.

OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model Pretraining

Yuran Wang; Siqiao Huang; Mingleyang Li +21 authors

OpenWAM turns world–action modeling into a modular design study, then instantiates OpenWAM-α: a video DiT coupled to a dedicated ActionDiT through mutual attention and joint denoising. Its strongest lesson is conditional: embodied pretraining improves scene transfer, but strong manipulation scores coexist with substantial visual-robustness failures. The evidence below separates controlled ablations, final-model benchmarks and physical execution.

Beyond Task Success: Stage-Wise Reliability of World Model Planning under Sensing Degradation

Geonmyeong Lee; Byoung-Tak Zhang

This diagnostic study follows synthetic sensing disturbances through an existing world-model planner's representation, action-conditioned prediction, candidate ranking, and executed task outcome. Paired OGBench evaluations show that a disturbance's relative impact can change between stages. Temporal information position matters even when aggregate representation shifts are similar. These diagnostics locate sensitivity; they neither establish a general predictor of task failure nor demonstrate successful sensing mitigation.

WM-Craftnet: World Synesthesia Model for Generalizable and Robust Dexterous In-Hand Manipulation

Jie Yin; Zeyuan Zhao; Xiaojing Tan +3 authors

WM-Craftnet learns a predictive visuotactile state that conditions a separate dexterous manipulation policy. Its main contribution is recurrent perception for executed control: noisy depth, touch, proprioception, and action history inform a latent state trained with clean-depth and other prediction targets. Hardware rotation improves substantially, but severe-disturbance recovery remains limited (e-rssm, e-policy, e-stress).

Learning to Use Imagination: Progress-Conditioned Future Utilization for World Action Models

Yijie Zhu; Zitong Yu; Wei Li +4 authors

ProWAM learns execution progress from recent action–observation feedback and recurrent memory, then uses that state to redistribute imagined-future attention inside a joint video–action policy. Its strongest evidence combines executed manipulation results with ablations separating progress estimation from future utilization.

How to Learn from What a Human Would Avoid? Intervention-Aware World Models with Real-World RL for Dexterous Manipulation

Jiaju Yin; Zhenhui Zhang; Lixin Xu +5 authors

WHIRL learns when a human would take over a dexterous robot and uses that prediction to steer residual-policy training. A frozen imitation prior supplies nominal behavior; an action-conditioned, one-step latent world model supports critic learning and actor risk shaping. Five real-robot tasks show higher autonomous success, but the evidence is restricted to one training seed, one operator and familiar workspace regions.

Learning Counterfactual World Models for Embodied Reasoning under Partial Observability

Todd Y. Zhou; Daniel Zhang

Counterfactual Latent World Models (CLWM) train action-conditioned imagined futures to preserve differences in intervention outcomes even when observations look alike. A recurrent world model supplies latent rollouts to model-predictive control; privileged outcome labels supervise an additional contrastive objective during training. Reported simulation gains reach 11.6 percentage points in navigation success. The evidence supports targeted representation training, while supervision-matched controls and audits of pretrained encoders remain open.

GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation

AgiBot Research Team

GE-Act 2.0 connects a compact visual representation, a one-step future generator, and a separate inverse dynamics model. Its central training problem is pairing generated futures with compatible recorded actions. KASO selects futures using the current action model before updating both modules. Real-robot experiments show broader manipulation capability as co-training data grows, with persistent weaknesses in fine manipulation and ordinal language; simulation comparisons use a separate adaptation protocol.

TacPAC: Tactile Prediction and Real-Time Action Correction in World-Action Models for Contact-Rich Manipulation

Zipei Ma; Xiaofei Wei; Junzhe Jiang +2 authors

TacPAC predicts future visual and tactile observations while planning an action chunk, then uses incoming touch to revise its unfinished actions against a fixed prediction-and-plan cache. On five physical manipulation tasks, it reports 64% average success versus 48% for T-Rex and 22% for its vision-only ablation. The central evidence is the controlled cache-access ablation, not visual prediction quality alone.

TourPhysics: Bringing Physics to World Models for Exploration and Manipulation from a Single Image

Xin Zhang; Yabo Chen; Zixuan Duan +4 authors

TourPhysics renders persistent camera tours and object manipulations from a single image plus a declared physical scene. A simulator fixes each trajectory before diffusion generates its appearance; accepted windows publish physical endpoints and visual memory together. The strongest evidence concerns prescribed motion and modest improvements on held-out revisits, with limited physical-event support and no real-world validation.

GIFT: Guided Intermediate Feature Training via Action-Oriented Structural Supervision for Robotic Manipulation

Yupeng Zheng; Xiang Li; Songen Gu +11 authors

GIFT trains robot-policy features to retain geometry, object–end-effector relations and instruction-relevant regions. The same auxiliary objectives improve three different action formulations while their default deployment uses no auxiliary predictions. Evidence is strongest for matched-baseline robustness gains, with substantial annotation requirements and uneven benefits across shifts.

Toward Unified Robot Learning: Bridging Representation, Vision-Language-Action, and World Models

Shaunak A. Mehta; Ananya Hazarika; Haochen Zhang +5 authors

This survey organizes robot learning around representations that encode the environment, VLA policies that generate actions, and world models that predict consequences. Its useful contribution is a vocabulary for tracing how information and feedback cross those interfaces. Integration can remain modular; the authors argue that its value should be demonstrated through improved behavior under uncertainty, distribution shift and temporal dependencies. A small original CALVIN video-prediction diagnostic illustrates hidden-object and collision failures, but supplies no quantitative proof that a unified architecture resolves them.

WISE: World-model-guided Imagination Scheduling for Efficient Post-training of Vision-Language-Action Models

Chenhao Zhang; Hanyu Zhao; Hang Cheng +2 authors

WISE post-trains a VLA action head by selectively imagining alternative behaviors at visually identified interaction states. A separate frozen world model predicts bounded futures; a frozen evaluator ranks them; updates supervise only the first action chunk from a real context. The strongest controlled evidence is the scheduling ablation: higher task success with substantially less imagination computation. The method, results, and reproduction boundaries below trace this conclusion to the primary text.

SV-WAM: An Efficient Surround-View World-Action Model for End-to-End Autonomous Driving

Jinyang Wang; Shiwei Li; Junjian Wang +12 authors

SV-WAM trains one shared transformer to denoise driving actions and future surround-view video, but blocks action attention to future-video tokens. Deployment retains six-camera history while generating only trajectories. A differentiable vehicle-footprint loss improves road compliance. The strongest reported result is 91.0 EPDMS on NAVSIMv2 navtest; the evidence concerns benchmark planning, with weaker hard-split performance and no demonstrated real-vehicle deployment.

Drive-HWM: Hierarchical World Models for Dynamic-Latent Guided Autonomous Driving

Zhaoxin Fan; Tianbao Zhang; Wenjun Wu +5 authors

Drive-HWM couples a periodically refreshed predictor of future motion latents with an observation-grounded autoregressive driving policy. Optical flow supervises the slow branch; next-frame RGB tokens supervise the fast branch during training. NAVSIM tables report stronger aggregate driving scores, but conflicting prose and incomplete implementation details limit precise reproduction.

Toward Physically Grounded JEPA World Models for Goal-Conditioned Robotic Planning

Muyuan Liu; Yue Huang; Zheng Liang +1 authors

SA+IDM trains an action-conditioned JEPA world model with two auxiliary heads: inverse dynamics recovers executed actions, while state alignment predicts measured physical state from consecutive image representations. Deployment uses only the encoder and latent predictor inside CEM planning. State alignment improves all four reported tasks over IDM alone, while the diagnostics challenge average temporal straightening as a sufficient representation-quality criterion (e2–e12).

Building Pretraining Data for World Models: An Unreal Engine-Based Pipeline for Action-Conditioned Video Generation

Haoyu Wang; Songchun Zhang; Haoran Li +3 authors

This paper documents the Unreal Engine synthetic-data component used in EchoWM: simulate character motion once, record controls and states, then replay the trajectory for high-quality multi-view rendering. Its contribution is a production system combining scene curation, cache locality and failure recovery. Reported video volume establishes production scale; downstream learning utility remains untested here. [e-scope, e-workflow, e-scale, e-limits]

Latent Energy Action Planning with World Models

Phu Pham and Aniket Bera

LEAP refines an action horizon through frozen LeWorldModel dynamics, combining latent-goal matching with a learned decoder’s terminal-state error. A trained proposal initializes search; projection bounds executed controls. Four-domain mean success rises from 77.5% to 94.8% against matched LeWM+CEM. The narrower energy ablation improves from 91.0% to 96.5%, separating the extra objective’s contribution from the complete planner change. Numerical goal descriptors and trustworthy learned rollouts remain prerequisites (e-energy, e-proposal, e-optimization, e-main, e-ablation, e-limits).

Spatially Aware World Action Model via Geometric Latent Diffusion

Gonzalez, Javier Alejandro Lopetegui; Pacaud, Paul; Schmid, Cordelia

To bridge this gap, we introduce a Spatially Aware World Action Model (SA-WAM), which repurposes a pretrained video model for joint action, RGB, and depth prediction, enabling 3D-aware world modeling and action prediction within a single diffusion backbone. SA-WAM achieves state-of-the-art results on the RoboCasa and LIBERO-Plus benchmarks, while simultaneously improving future-state predictions.

Latent Action as Intention Enables Efficient Future Imagination for World Action Models

Li, Xiang; Zheng, Yupeng; Gu, Songen +11 authors

To bridge this gap, we introduce LAWA, a WAM architecture that uses compact latent actions as an operational representation of future intentions, enabling efficient test-time future imagination without generating future observations. On RoboCasa, LAWA achieves state-of-the-art average success rates of 65.6% and 80.8% in the few-shot and full data settings, improving over the matched Fast-WAM baseline by 9.6 and 4.5 points, respectively.

DELE-w0.5: Inferring Action from Future Latent State for Robotic Manipulation

Lei, Fenghao; Huang, Zhixiong; Yang, Long +5 authors

In this paper, we propose DELE-w0.5, which infers robot actions from predicted future states without relying on video generation. Across 640 real-robot trials on four long-horizon manipulation tasks, our DELE-w0.5 achieves the best performance among all compared policies, attaining 62.5% overall full-task success and 81.3% macro ordered-stage progress.

Flex-ππ: A Multi-Stream World-Action Model with Compute Flexibility

Yan, Ge; Liu, Jinghao; Fan, Yuzhi +4 authors

World-action models (WAMs) predict the future to act better, but nearly all of them predict only RGB latents, trained purely for pixel reconstruction, with no explicit signal for the 3D geometry or object semantics manipulation needs. We find a surprising free lunch: the same frozen video-generation VAE that encodes RGB also encodes 3D pointmaps almost losslessly, with no pointmap-specific training at all.

Hydra: A Navigation World Action Model with Discrete Latent Planning and Continuous Flow-Matching Execution

Nazeri, Mohammad; Card, Alexandyr; Huber, Samira +6 authors

In this paper, we present Hydra, a discrete World Action Model that closes this gap by moving the planner, both the sampler and the evaluator, inside the model. Evaluated on two physical robotic platforms, Hydra outperforms state-of-the-art world models in goal-directed planning, while matching or exceeding the closed-loop execution capabilities of leading reactive foundation policies.

Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models

Yang, Senqiao; Wang, Chengyao; Chen, Yuxin +13 authors

We propose VLAct, a VLA-oriented VLM backbone trained on broad, heterogeneous, multi-embodiment robot data before task-specific fine-tuning. On RoboDojo, VLAct ranks sixth among all policies by success rate and outperforms all explicitly designated world-action model (WAM) entries on both metrics.

Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization

Zhou, Jiaming; Zhang, Qihang; Xu, Gangwei +9 authors

To achieve cross-task generalization, we bring this paradigm to robotic manipulation, and argue that the natural task specification for manipulation is a human video: unlike language, it provides rich visual cues about the intended task evolution. We present Zero-WAM, a causal video-action model that executes unseen tasks by following in-context human video guidance.

GlanceWAM: Sparse Test-Time Imagination for World-Action Models

Wang, Linhan; An, Zijian; Zhang, Mingyuan +7 authors

We show that visual imagination achieves both real-time inference and superior success rates when generated asynchronously off the critical path and consumed directly in latent space. We introduce GlanceWAM, which decouples imagination from control within a single video DiT: an asynchronous proposer glances ahead on a slow clock to imagine a single lookahead frame seconds into the future in the background, while an action head decodes action chunks at control rate (48 ms) purely in latent space without blocking.

LD4WAM: Learning Latent Dynamics from Human Videos for World Action Models

Shen, Zhenhao; Liang, Jiaqi; Lu, Jasper +11 authors

Human video is playing an increasingly central role in training World Action Models (WAMs), owing to its diversity and low collection cost relative to teleoperated robot data. However, most WAMs learn from such video only by predicting pixel-level future frames, giving dynamics that are not directly actionable, whereas motion retargeting recovers directly actionable actions but leaves a large visual gap across embodiments.

ForeTime-VLA: Causal Future-Token Distillation from a World Action Model for Conveyor-Belt Manipulation

Ma, Siyuan; Zhang, Yutian; Zhang, Boshi +4 authors

We introduce ForeTime-VLA, a dense pi0.5 policy that distills a future-aware, action-equivalent representation from a frozen Fast-WAM-derived teacher while remaining causal at inference. In quantitative real-robot evaluation, ForeTime-VLA achieves 81.1% stationary and 58.9% slow-moving grasp success, exceeding the next-best reference by 12.2 and 22.2 percentage points, respectively.

On the Capability Separation Between World-Model Policy Learning and Imitated World-Action Models

Yu, Yang

World-action models predict a future outcome and then infer an associated action. Although this factorization can improve representation learning and data efficiency, it is unclear whether it provides stronger control capability than direct behavior cloning when both are trained from the same observational demonstrations.

WA-JEPA: Rethinking the Video JEPA Paradigm for World-Action Modeling in Autonomous Driving

Wang, Xinlin; Xiang, Yujiao; Zhou, Yuheng +11 authors

Video Joint Embedding Predictive Architecture (V-JEPA) learns powerful spatiotemporal representations from video through self-supervised latent feature prediction. However, V-JEPA is built around random-mask completion and deterministic regression, making it fundamentally ill-suited for autonomous driving planning that demands future-directed prediction tightly coupled with action.

Towards Surgical World-Action Modeling: A Preliminary Joint Visual-Trajectory Forecasting for Surgical Motion Planning

Huang, Weiliang; Liu, Huanrong; Zhang, Bob +5 authors

To bridge this gap, we present a preliminary joint visual-trajectory world-action model that simultaneously forecasts future visual states and instrument trajectories from historical surgical observations. The chunked strategy consistently outperforms direct one-shot prediction across all evaluated horizons, improving first-segment PSNR from 18.86 to 23.11 dB and reducing ADE from 45.77 to 22.22 pixels.

What Matters for Latent Actions in Robot Learning

Xizhou Bu; Qingda Hu; Lei Zhou +13 authors

This empirical study asks which video-derived latent actions improve robot policies. It compares modeling paradigms, regularization, and action integration under a shared training framework. Raw-frame LAPO remains competitive, but preferred dimensionality and action heads depend on the evaluation. Its strongest deployment evidence is improved Franka manipulation after latent-action tuning of a VLM backbone; the forward predictor supplies training supervision rather than an inference-time planner.

RISE: Adaptive Imagination for World Action Models

Lu, Hongbo; Yao, Liang; He, Chenghao +5 authors

We propose RISE (\textbf{R}efining \textbf{I}magination through \textbf{SE}lective Rollout), a system-level adaptive imagination framework that makes sequential \textsc{Roll}/\textsc{Stop} decisions according to the expected planning benefit of continued rollout. Experiments on NAVSIM and nuScenes show that RISE achieves the best overall planning performance while reducing unnecessary rollout, with additional transfer results supporting its plug-in generality across WAM architectures.

HiTac-WAM: A Hierarchical Tactile World Action Model for Contact-Rich Robot Manipulation

Xue, Chao; Zhang, Chaofan; Ma, Wenxuan +3 authors

We present HiTac-WAM, a hierarchical tactile world action model that forecasts a sequence of future tactile states for each candidate action chunk before execution. HiTac-WAM achieves a mean contact F1 of 0.921; under matched training budgets, the directed hierarchy reduces 3D displacement L2 error by 17.6% relative to the deformation-only predictor and improves slip AUPRC by 60.4% relative to the slip-only predictor.

BrainWAM: Action-Space Coordination of Semantic Priors and Predictive Dynamics for Autonomous Driving

Zhan, Bing; Shang, Shuyao; Lu, Shuo +6 authors

Inspired by neuroscience evidence that complex behavior arises from coordination among functionally specialized systems, we propose BrainWAM, a structured action-space coordination framework that converts semantic reasoning and predictive world modeling into two specialized action-oriented pathways, and aligns them at the level of compact action representations. BrainWAM reaches state-of-the-art performance on both NAVSIM v1 (89.5 PDMS) and NAVSIM v2 (89.6 EPDMS), consistently outperforming VLA-only or WAM-only methods, highlighting BrainWAM as a practical and promising direction for autonomous driving systems.

StageWAM: Joint-Embedding Stage Prediction for World-Action Models in Robot Manipulation

Liu, Xiao; Yang, Yuguang; Wang, Xi +6 authors

We introduce StageWAM, which augments a Motus-based World Action Model (WAM) with Stage-JEPA, a goal-conditioned Joint-Embedding Predictive Architecture (JEPA) predictor. Across 50 RoboTwin 2.0 tasks in clean and randomized environments, StageWAM achieves 90.25% overall success and reduces the mean number of execution steps in successful rollouts by 5.97% relative to the strongest baseline.

PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud

Chenghua Wang; Daliang Xu; Dongqi Cai +23 authors

To unify them, we build PhyAI, a Physical AI inference engine with a single runtime that keeps architecture-specific conditioning, solver, cache, and output logic in model adapters while sharing graph execution, kernels, memory management, and parallel services. PhyAI achieves 1.40x-4.65x speedups over the official implementations of pi0, pi0.5, GR00T N1.7, and MiniCPM-Robot.

4D-WAM: Infusing Spatiotemporal Awareness into World Action Models through Trajectory Fields

Yang, Lishan; Song, Wenxuan; Wang, Xi +14 authors

In this work, we propose 4D-WAM, a model-agnostic training strategy that injects spatiotemporal knowledge from 3D trajectory fields into WAMs through representation alignment. To this end, we introduce two complementary objectives: 1) motion alignment, which aligns temporal feature variations across adjacent frames and encourages the model to build local 4D awareness during training, and 2) destination alignment, which guides the model to infer the final destination from the source frame by minimizing the gap between their attention-like similarity distributions.

World Action Models in Real Time: An Empirical Study of Smooth Execution via Asynchronous Deployment

Team, Motubrain

We present an empirical study of asynchronous deployment strategies that overlap model inference with action execution to enable responsive and smooth control. In contrast, prefix-conditioned generation achieves the best overall balance between task performance, execution speed, and trajectory smoothness by learning consistent action continuations during training.

FACT: Failure-Aware Causal Training for World-Action Models

Peng, Quanquan; Liang, Yutong; Yan, Rui +2 authors

We introduce FACT, a causal World-Action Model that predicts future video and task progress conditioned on the executed action. Extensive experiments on simulation and real-world bimanual manipulation tasks show that FACT outperforms many existing baselines, improves as failure data are incorporated into training, and reduces success-biased future hallucination under bad actions.

4D-WAM: 4D Consistent World Modeling for Autonomous Driving

Fu, Jiacheng; Yuan, Yibo; Tian, Meng +8 authors

To alleviate this issue, we present 4D-WAM, a model that leverages geometric foundation models for training-time supervision to enable 4D consistent world modeling. Extensive experiments demonstrate that 4D-WAM effectively models 4D consistent scene evolution and achieves state-of-the-art performance on challenging NAVSIM-v1 and NAVSIM-v2 benchmarks.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling

Lin, Yihan; He, Jiawei; Bao, Shifeng +6 authors

We introduce JEPA-WAM, a latent WAM built in a pretrained V-JEPA space, which couples latent transition prediction with continuous action generation through a shared predictor. On LIBERO-Plus, JEPA-WAM achieves 79.2%, the best result without large-scale robot-policy pretraining, while its pretrained π0.5π_{0.5} instantiation reaches 86.3%, achieving the best overall performance.

Vid2WAM: Distilling Video Diffusion Priors into World Action Models

Qiu, Chenhao; Wang, Ruixiang; Zhao, Runyi +7 authors

In this paper, we propose Vid2WAM, an offline distillation framework that transfers visual diffusion priors from a large video foundation model into a compact WAM student. To robustly integrate synthetic and real supervision, we introduce source-aware residual action adaptation that learns source-specific corrections around a shared action backbone and mitigates interference from noisy pseudo-actions.

ω-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation

Li, Zhe; Zhang, Zhenzhe; Wei, Yangyang +8 authors

We present ωω-0, a latent predictive whole-body world-action model for real-world humanoid concurrent loco-manipulation. Real-world experiments on 11 household tasks demonstrate that a single ωω-0 model can produce smooth manipulate-while-moving behaviors and consistently outperform representative imitation learning, VLA, humanoid, and WAM baselines.

Selective Cross-View Consistency for World Action Models: Held-Out Viewpoint Robustness Without Test-Time Camera Information

Huang, Bingqi; Wei, Bingchuan; Cai, Yingkai +1 authors

We introduce a carve-and-hold-out evaluation protocol on the LIBERO-Plus camera track that separates a distribution-matched ceiling from genuine interpolation and extrapolation to held-out viewpoints, with a matched pair-trained control isolating the effect of the consistency term from pair exposure.

Faster-WAM: Efficient Inference-Time Future Conditioning for Robust World Action Models

Zhao, Weiheng; Jiang, Haoyi; Shi, Xin +5 authors

Specifically, we propose SparseMoT to replace ubiquitous layer-wise fusion with selective video-action interaction at a compact subset of network stages, and Interval KV-Fusion to aggregate multi-depth future representations without increasing attention complexity. Experiments demonstrate that Faster-WAM achieves a substantially better performance-efficiency trade-off than existing WAMs.

UniNav: A Unified World-Action Diffusion Model for Visual Navigation

Zhou, Changqing; Luo, Yueru; Jiang, Zeyu +1 authors

We present UniNav, a unified world-action model that generates future visual observations and continuous waypoint trajectories through a single diffusion process. Based on this unified framework, we introduce two variants: UniNav-Full jointly predicts interpretable future observations and their corresponding trajectories, while UniNav-Fast removes future-image tokens at inference for efficient trajectory prediction.

Faster-WAM: Do World Action Models Need Deep Action Modules?

Ma, Liheng; Yang, Rui Heng; Zhang, Zhanguang +4 authors

To address this limitation, we introduce Dock of Transformer (DoT), a video-centric design principle that treats a pretrained video Transformer as a representation hub and connects lightweight output-heads through docking interfaces. Without additional embodied pretraining, Faster-WAM achieves competitive performance on LIBERO and RoboTwin 2.0 while demonstrating strong out-of-distribution generalization on LIBERO-Plus.

SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space

Zhao, Ruiteng; Zhang, Zhengshen; Su, Yue +6 authors

We propose SG-WAM, a self-guided framework that learns geometry-aware action-conditioned dynamics directly in the policy-derived representation space. Built on a 0.9B model without large-scale embodied pretraining, SG-WAM achieves 98.5% average success on LIBERO and 73% on LIBERO-Plus, while outperforming strong baselines in both in-distribution and out-of-distribution real-world evaluations.

SelfWAM: A Self-Grounded Unified World Action Model for Fast Robot Control

Pan, Bikang; Liu, Fan; Lu, Haotao +2 authors

We introduce SelfWAM, a unified self-grounded WAM built on a modality-specialized Mixture-of-Transformers (MoT) architecture that jointly predicts actions, action-conditioned future RGB frames, and robot self-masks, thereby grounding future prediction in the robot's visible body and its action-induced motion.

Disentangling Visuo-Tactile Foresight: Oracle-Guided Interface Discovery for World Action Models

Yao, Zihang; Ding, Chaoyue; Yu, Yingying

To make this interface independently studyable, we introduce Oracle Visuo-Tactile Foresight (OVTF), a controlled framework that supplies paired RGB and tactile futures from successful trajectories verified in simulation. Within OVTF, we propose Asymmetric Phase-Local Future Memory (AFM), in which visual memory reads future vision, each tactile memory jointly attends to its own tactile stream and phase-aligned future vision, and cross-tactile access is blocked.

Paired Exact-Reset Evaluation of a Prediction-Derived Medium-to-Full World-Model Cascade

Pastor, Malo de

Existing adaptive-inference and world-action-model systems use cheap-stage outputs or predicted futures to allocate additional computation. We study a narrower question: under paired exact-reset physical outcomes, can a Medium-derived interface predict when switching to a separately frozen Full predictor improves task-specific decision loss enough to justify sequential overhead?

ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts

Wang, Mingxin; Hu, Bin; Qian, Bin +12 authors

Rather than correcting the predicted futures, we propose Semantic-Temporal WAM (ST-WAM) to improve action robustness by using DINOv3 as a shared semantic representation for future prediction and history retrieval while retaining fine-grained VAE dynamics. It achieves 98.7% on LIBERO and 92.8% on RoboTwin 2.0; more importantly, compared with Fast-WAM, it improves zero-shot LIBERO-Plus performance by 21.3 percentage points and more than doubles real-world success under visual shifts from 25.8% to 61.5%.

EgoGenesis: Egocentric World-Action Modeling with Online Anchored Projective Memory and Action-3D RoPE

Yan, Zexuan; Wu, Yuzhou; Ma, Yue +9 authors

We present \method, an egocentric world-action simulator that synthesizes controllable, high-quality manipulation videos to expand scarce real-world training data. \method{} builds on a pretrained video generation prior and introduces two geometry-aware conditioning mechanisms.

LeapBot-WA: World-Anchor Action Models via Predictive Latent Alignments

Liu, Pei; Zheng, Nan; Zhang, Lang +8 authors

In this paper, we propose LeapBot-WA, which establishes a novel Predictive-Latent paradigm for WAMs by operationalizing the Joint-Embedding Predictive Architecture (JEPA) as a World-Anchor. To bridge the modality gap between non-Gaussian predictive features and diffusion priors, we introduce the Isotropic Semantic Autoencoder (ISAE), which reshapes the anchor's latent space into a diffusion-friendly manifold to prevent off-manifold drift.

HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone

AI, Simple; :; Wei, Yuteng +16 authors

We present HiFi-UMI, a portable UMI data-production system co-designed for trajectory accuracy, inter-gripper relative pose, synchronization, and field of view: head-mounted offline stereo-inertial SLAM, native rather than reconstructed relative pose, a shared microsecond GPIO trigger, and two wide-angle cameras per hand covering ~200 degrees.

GeoWorldAD: Geometry World Action Model for Autonomous Driving

Zhang, Songyan; Tian, Jinyuan; Li, Hanbing +9 authors

In this work, we propose GeoWorldAD, a geometry world action model that grounds trajectory planning in ego-aligned 3D space and anticipates short-horizon scene evolution with latent future geometry tokens. Experiments on NAVSIM v1 and v2 demonstrate state-of-the-art performance, highlighting the effectiveness of explicit 3D geometry grounding and future geometry world modeling for safe and efficient autonomous driving.

NVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulation

NVIDIA; Aarti Basant; Amlan Kar +31 authors

To overcome these limitations, we introduce OmniDreams, a foundation generative world model mid- and post-trained from the Cosmos diffusion model to autoregressively generate action-conditioned videos in real time. We additionally show preliminary results indicating that a world-action model (WAM) post-trained from OmniDreams achieves strong performance on the Physical AI Autonomous Vehicles NuRec dataset, surpassing the VLA-based Alpamayo 1.5 research policy model while using only 1/5 the total parameters.

From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence

Liang, Yuanzhi; Zhan, Xufeng; Huang, Haibin +2 authors

Building on this analysis, we propose a co-evolution roadmap for physical intelligence centered on the \emph{embodied brain}, a long-term model target for integrating multimodal context, comparing candidate interventions, and issuing state-transition or capability requests rather than direct actuator commands.

Artificial Foveated Perception for Mitigating Shortcut Learning in Robotic Foundation Models

Sun, Xiatao; Zhuang, Yuan; Negrete, Mateo Sanchez Lopez +9 authors

We propose Artificial Foveated Perception (AFP), a lightweight, policy-agnostic module that takes the same vision and language inputs as Vision-Language-Action and World Action Model pipelines and predicts task-conditioned masks over relevant objects, the robot, and other action-critical regions. We evaluate AFP across state-of-the-art robotic foundation models and show that foveated perception reduces fine-tuning time, suppresses overfitting, and improves generalization under environmental perturbations.

FlowDAgger: Human-in-the-Loop Adaptation of Generative Robot Policies in Latent Space

Murray, Michael; Chen, Daphne; Bagaria, Simran +7 authors

We present FlowDAgger, a sample- and compute-efficient method for adapting frozen generative robot policies from human interventions in latent space. FlowDAgger outperforms supervised fine-tuning and latent-space RL baselines and preserves pretrained skills on held-out tasks, offering a practical path for adapting robot foundation models in the real world.

Understanding and Mitigating the Video-Action Generalization Gap via Temporal Ratio

Mishra, Utkarsh A.; Chen, Yongxin; Xu, Danfei +3 authors

To explain this behavior, we introduce the Temporal Ratio (TR), an attention-based measure of how strongly the action head relies on future latent rollouts relative to the anchored current frame. Finally, based on these findings, we propose an inference-time adaptive guidance method, which exploits this intrinsic feature attention pattern to dynamically amplify compositional video conditioning signals precisely when the policy relies on future rollouts.

WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time

Feng, Yusen; Han, Bingchen; Lyu, Jiangran +13 authors

We present WAM-TTT, a test-time training framework for steering world action models from raw human videos. To make this memory useful for control, we introduce a meta-training stage that aligns human demonstrations with robot behaviors using paired human-robot data and a key--value memory reconstruction objective.

WAM4D: Fast 4D World Action Model via Spatial Register Tokens

Li, Ying; Wei, Xiaobao; Cao, Jiajun +10 authors

To address the trade-off, we present WAM4D, a fast 4D world action model that uses lightweight spatial register tokens as training-time future-depth readouts to transfer pretrained geometric priors into a causal video-action transformer, then removes the register branch for lightweight action inference. Comprehensive experiments on RoboTwin 2.0 and challenging real-world manipulation tasks show that WAM4D improves spatial consistency and achieves competitive action prediction while maintaining efficient inference.

Learning 4D Geometric Priors for Inference-Efficient World Action Models

Zhang, Jianjun; Zhu, Jian; Su, Taiyi +4 authors

We propose MECo-WAM, a Multi-Expert Co-Training World Action Model that injects action-relevant 4D geometric priors into video-action representations while preserving the original lightweight inference graph. To transfer geometric knowledge into the deployed video-action pathway, we introduce decayed 4D read-mask attention, which provides restricted current-frame geometric guidance early in training and progressively removes this dependency.

UNIVERSE: Unified Video Action Models for Autonomous Driving with Flexible Mask-Modulated Modality Generation

Liu, Mengmeng; Zhang, Diankun; Liu, Jiuming +7 authors

We propose UNIVERSE, a unified video-action model built upon a single mask-modulated Diffusion Transformer. To ensure causal validity and efficient deployment, we introduce a Modality-Decoupling Visibility Mask, which shares historical context across modalities while blocking mutual attention between future video and trajectory tokens.

DSWAM: A Dual-System World Action Foundation Model for Fine-Grained Robot Manipulation

Zhu, Jian; Zhang, Jianjun; Su, Taiyi +10 authors

To address both the decomposition gap and the need for a controlled WAM-VLA comparison, we introduce DSWAM, a Dual-System World Action Foundation Model for fine-grained robot manipulation. To provide a fair real-robot comparison with VLA policies, we build and evaluate DSWAM under the DeMaVLA real-world deformable manipulation setting with matched robot platform, pretraining data, post-training data, and evaluation criteria.

ABot-M0.5: Unified Mobility-and-Manipulation World Action Model

Chen, Ronghan; Yang, Yandan; Tang, Zuojin +18 authors

We propose ABot-M0.5, a new WAM built on the insight that mobile manipulation requires alignment at three levels: temporal granularity, action space, and train-test consistency. To align temporal granularity, we introduce intermediate latent actions that capture local visual state transitions and serve as an bridging action space between video latents and embodiment-specific controls.

HALO-WA: Hybrid-Attention Latent-Guided Online Reinforcement Learning for World-Action Models

Ye, Angen; Ke, Weijie; Wang, Xiaofeng +7 authors

We propose HALO-WA, a hybrid-attention latent-guided online reinforcement learning (RL) framework for WA models, which leverages latent features and action priors from the WA generation process through a lightweight actor-critic adapter to enable fast online adaptation to real deployment errors.

Kairos: A Regret-Aware Native World-Action Model Stack for Physical AI

Team, Kairos; Wang, Fei; You, Shan +21 authors

We introduce \textbf{Kairos}, a regret-aware native world-action model stack for Physical AI. Experiments on embodied world-model benchmarks, world-action benchmarks, long-horizon generation, and inference-efficiency evaluation show that Kairos achieves superior performance while offering a favorable efficiency to capability trade-off.

VT-WAM: Visual-Tactile World Action Model for Contact-Rich Manipulation

Tian, Shuai; Zheng, Yupeng; Zheng, Yuhang +7 authors

In this paper, we introduce VT-WAM, a Visual-Tactile World Action Model that jointly learns future visual prediction, tactile deformation prediction, and action prediction within a unified flow matching framework. Across six real-world contact-rich manipulation tasks, VT-WAM achieves a 71.67% average success rate, outperforming Fast-WAM by 26.67% and OmniVTLA by 35.84%.

FutureNav: Unified World-Action Modeling for Vision-and-Language Navigation

Zhang, Lingfeng; Gong, Zeying; Hao, Xiaoshuai +7 authors

We introduce FutureNav, a VLM-based unified world-action modeling framework for vision-and-language navigation. Extensive experiments show that, with only a 4B-scale backbone, FutureNav achieves state-of-the-art performance on multiple VLN benchmarks and substantially outperforms prior VLN methods, paving the way toward future world-action models for VLN.

Pondering the Way: Spatial-perceiving World Action Model for Embodied Navigation

Chen, Hong; Liu, Daqi; Zhang, Zehan +10 authors

To address these issues, we propose SWAM (Spatial-perceiving World Action Model), a task-centric joint observation-action generation framework. Extensive experiments show that SWAM significantly outperforms state-of-the-art two-stage planners in success rate, trajectory accuracy, and inference efficiency, while demonstrating robust zero-shot generalization to unseen environments.

World Action Models Enable Continual Imitation Learning with Recurrent Generative Replays

Govind, Manish Kumar; Reilly, Dominick; Patel, Smit +2 authors

We build on this generative capability to propose Recurrent Generative Replay (REGEN), a continual imitation learning framework that synthesizes pseudo-replay trajectories, enabling a robot policy to rehearse previously learned tasks without storing their original human demonstrations.

PAIWorld: A 3D-Consistent World Foundation Model for Robotic Manipulation

Yuhang Huang; Xuan Lv; Junyan Xu +25 authors

To address this, we present PAIWorld, a framework that augments diffusion-transformer world models via three core components: (1) Geometry-Aware Cross-View Attention blocks that establish an explicit pathway across views, (2) Geometric Rotary Position Embedding that encodes camera ray directions and extrinsic poses into the attention mechanism, and (3) Latent 3D-REPA, which distills 3D-aware features from frozen 3D foundation models to ensure 3D consistency. Built upon a DiT-based world foundation model, PAIWorld achieves state-of-the-art multi-view 3D consistency on robotic manipulation benchmarks, ranking 1st on the WorldArena leaderboard and 2nd on the AgiBot-Challenge2026 leaderboard, wh…

Cosmos 3: Omnimodal World Models for Physical AI

NVIDIA; Aditi; Niket Agarwal +292 authors

We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-transformers architecture. Our evaluation demonstrates that Cosmos 3 establishes a new state-of-the-art across a diverse suite of understanding and generation tasks, demonstrating omnimodal world models as scalable, general-purpose backbones for embodied agents.

dVLA-RL: Reinforcement Learning over Denoising Trajectories for Discrete Diffusion Vision-Language-Action Models

Wu, Yuhao; Liu, Yitian; Shen, Weijie +13 authors

To solve this problem, we propose \textbf{dVLA-RL}, shifting the learning objective from the marginal action probability to the joint probability of the sampled generation path. Leveraging this intrinsic fexibility, we introduce a unified step scheduling approach for complex multi-task learning, tailoring denoising steps to specific task complexities to maximize both success rates and computational effciency.

MV-WAM: Manifold-Aware World Action Model with Value Augmentation

Chen, Jintao; Jia, Peidong; Wuwu, Qingpo +13 authors

Recent world action models achieve strong in-domain performance, yet their gains do not extend proportionally to out-of-distribution scenarios. To address this, we propose MV-WAM, a novel end-to-end framework that jointly models visual prediction, action generation, and value estimation designed to effectively leverage video priors during both training and inference for enhanced action generalization.

MemoryWAM: Efficient World Action Modeling with Persistent Memory

Yang, Sizhe; Mu, Juncheng; Wei, Tianming +8 authors

To address this challenge, we introduce MemoryWAM, a world action model with efficient persistent memory. Across long-horizon, memory-dependent manipulation tasks in both simulation and the real world, MemoryWAM outperforms strong vision-language-action (VLA) and WAM baselines while maintaining favorable computational efficiency.

WAM-RL: World-Action Model Reinforcement Learning with Reconstruction Rewards and Online Video SFT

Qian, Zezhong; Chi, Xiaowei; Qi, Yu +3 authors

To address these limitations, we propose WAM-RL, a reinforcement learning framework that enables joint optimization of the world model and the action model through online interaction with the environment. By allowing the two components to co-evolve, our approach enhances fine-grained control and adaptability.

Metis: A Generalizable and Efficient World-Action Model for Autonomous Driving and Urban Navigation

Li, Jingyu; Liu, Zhe; Hu, Dongnan +10 authors

To address both issues, we propose Metis, an end-to-end WAM framework that decouples video generation and action prediction. To enhance efficiency, we introduce an asymmetric attention mask that enables joint training of both experts while allowing the action model to bypass explicit video generation during inference.

LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies

Chen, Jialei; Wang, Kai; Chen, Kang +9 authors

We present LaWAM, a Latent World Action Model that exposes predictive dynamics to robot policies through compact latent visual subgoals instead of reconstructed future video. LaWAM achieves state-of-the-art or competitive success rates (SRs) across LIBERO (98.6% SR), RoboTwin (91.22% SR), and real-world manipulation tasks while retaining low-latency inference.

WAM-Nav: Asymmetric Latent World-Action Modeling for Unified Visual Navigation

Yang, Ning; Huang, Yan; Peng, Kaiwen +9 authors

To address these limitations, we propose WAM-Nav, a Latent World-Action Model for embodied visual navigation that jointly learns action generation and latent visual foresight, enabling more robust and foresighted navigation decisions without compromising inference efficiency. To further encourage smooth and consistent trajectory generation, we introduce a dual-stream contextual conditioning mechanism that integrates episode-level ego-motion history with sequential visual observations.

NavWAM: A Navigation World Action Model for Goal-Conditioned Visual Navigation

Azuma, Daichi; Miyanishi, Taiki; Sakamoto, Koya +6 authors

We propose Navigation World Action Model (NavWAM), a diffusion-transformer policy that turns navigation world-model prediction into executable action by representing future observations, goal-progress values, and action chunks in a shared latent sequence. We build NavWAM through simulation pretraining and real-robot adaptation, and evaluate it on image-goal navigation against planning-based world models and a representative direct navigation policy.

Diffusion Transformer World-Action Model for AV Scene Prediction

Sharifullin, Ruslan; Jiang, Benjamin; Chew, Kai Xi

Action-conditioned world models let an autonomous vehicle predict future camera scenes from its own planned controls, enabling planning and simulation without real-world rollouts, but at compact, trainable scale the futures are ambiguous and the field's standard distortion metrics actively mislead: they reward a blurry regression mean over a realistic prediction. We confront this with a compact latent world model that, given the present front-camera latent and a sequence of ego-actions, predicts future scene latents a frozen decoder renders to 256×256256 \times 256 frames up to 8 seconds ahead, evaluated on 150 held-out nuScenes scenes.

World Pilot: Steering Vision-Language-Action Models with World-Action Priors

Lin, Zefu; Cui, Rongxu; Xu, Junjia +4 authors

We present World Pilot, a VLA framework that augments the policy with priors from a World-Action Model (WAM), routed into the decision chain through two complementary pathways. World Pilot attains a state-of-the-art Total success rate of 84.7% on the LIBERO-Plus zero-shot OOD benchmark and the highest success rate on every real-robot setting across four manipulation tasks, with the largest margins under shifts in viewpoint, geometry, deformable state, and pose.

Next Forcing: Causal World Modeling with Multi-Chunk Prediction

Xu, Gangwei; Zhang, Qihang; Zhou, Jiaming +4 authors

In this paper, we present Next Forcing, a multi-chunk prediction (MCP) framework for causal world modeling that enables faster training, higher accuracy, and accelerated inference. During training, the MCP modules significantly accelerate convergence and improve converged accuracy, especially at high frame rates: at 50 fps, Next Forcing achieves a 93.1% relative improvement over LingBot-VA at 5k training steps and 2.3x faster convergence, and establishes new state-of-the-art results on the RoboTwin benchmark (94.1/93.5% on Clean/Random).

HiMem-WAM: Hierarchical Memory-Gated World Action Models for Robotic Manipulation

Sun, Xiaoquan; Zhang, Ruijian; Cao, Chen +12 authors

To address this, we present HiMem-WAM, a Hierarchical Memory-Gated WAM that integrates motion-centric latent actions, high-level skill latents, and boundary-triggered memory updates. Specifically, we develop a hierarchical latent action framework that jointly learns low-level motion and high-level skill latents, providing structured temporal abstraction.

AHA-WAM:Asynchronous Horizon-Adaptive World-Action Modeling with Observation-Guided Context Routing

Cai, Jisong; Ling, Long; Chu, Shiwei +10 authors

Therefore, we propose AHA-WAM, an Asynchronous Horizon-Adaptive World-Action Model built on a dual Diffusion Transformer (DiT) architecture that reorganizes world-action modeling around this temporal asymmetry. To support asynchronous execution, we introduce horizon-adaptive offset training and Observation-Guided Video-Context Routing (OVCR), which together let the action expert exploit long-horizon world context while remaining responsive to real-time execution state without rerunning the video DiT.

MotionWAM: Towards Foundation World Action Models for Real-Time Humanoid Loco-Manipulation

Zheng, Jia; Ma, Teli; Fan, Yudong +3 authors

We present MotionWAM, a real-time WAM that drives autonomous humanoid loco-manipulation from a single egocentric camera by conditioning the policy on the intermediate denoising features of a video world model. On nine real-world Unitree G1 tasks, MotionWAM runs in real time, substantially outperforms Vision-Language-Action (VLA) baselines fine-tuned on the same demonstrations by over 30% in overall success rate, and executes task-driven foot interaction that decoupled upper-lower policies cannot reach.

C3^3ache: Accelerating World Action Models with Cross Inference Chunk Cache

Zhao, Weisen; Nguyen, Lam; Lu, Zhicong +1 authors

We introduce C3^3ache, a training-free method that caches and reuses these residuals across inference chunks at the same denoising step. Experiments on benchmarks with a Fast-WAM backbone show that C3^3ache achieves up to a 2.5×2.5\times speedup in total wall-clock inference time, with negligible degradation in task success rate.

Flash-WAM: Modality-Aware Distillation for World Action Models

Akbari, Arman; Zhang, Ci; Akbari, Arash +6 authors

We introduce \textbf{Flash-WAM}, a modality-aware step-distillation framework inspired by consistency distillation that selects the consistency function for each modality to match its noise regime: a linear-gradient-scaling parametrization for the action stream's low-noise regime, paired with a variance-preserving parametrization for the video stream's high-noise regime, grounded in a structural analysis of the consistency-function family that characterizes the achievable gradient scaling under the consistency boundary condition.

GeoSem-WAM: Geometry- and Semantic-Aware World Action Models

Ma, Fulong; Peng, Daojie; Yue, Wenjun +4 authors

To address this, we propose a structured world modeling framework that enhances latent representations through geometric and semantic supervision. Alongside future RGB prediction, our model introduces two auxiliary prediction branches for future geometry and semantic representations, enabling it to jointly capture scene dynamics, spatial geometry, and semantic context within a unified latent space.

DriveWAM: Video Generative Priors Enable Scalable World-Action Modeling for Autonomous Driving

Shi, Chen; Xu, Jinrui; Shi, Shaoshuai +3 authors

We present DriveWAM, a driving world-action model that adapts a pretrained video diffusion transformer into an autoregressive video-action policy. To incorporate high-level scene understanding, we introduce scene-evolving driving guidance, where a frozen VLM produces chunk-specific semantic intent to guide video-action generation.

OASIS: Observation-Action Space Alignment via SE(3) Trajectory Prediction for Robotic Manipulation

Chen, Xinzhe; Ren, Sihua; Huang, Liqi +5 authors

We propose OASIS, a visuomotor policy that aligns the intermediate representation with the action space via SE(3)SE(3) end-effector trajectory prediction. Across simulation and real-world experiments, OASIS outperforms VLA and WAM baselines in success rate and out-of-distribution generalization.

World Action Models: The Next Frontier in Embodied AI

Wang, Siyin; Shi, Junhao; Fu, Zhaoyang +11 authors

Vision-Language-Action (VLA) models have achieved strong semantic generalization for embodied policy learning, yet they learn reactive observation-to-action mappings without explicitly modeling how the physical world evolves under intervention. A growing body of work addresses this limitation by integrating world models, predictive models of environment dynamics, into the action generation pipeline.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models

Feng, Qiuxuan; Yu, Jiale; Liu, Jiaming +8 authors

Motivated by these findings, we propose HarmoWAM, an end-to-end WAM that fully leverages a world model to unify predictive and reactive control, enabling both generalizable transit and precise manipulation. Notably, HarmoWAM achieves strong zero-shot generalization across these scenarios, significantly outperforming prior state-of-the-art VLA models and WAMs by margins of 33% and 29%, respectively.

When to Trust Imagination: Adaptive Action Execution for World Action Models

Wang, Rui; Zhang, Yue; Lin, Jiehong +4 authors

In this work, we formulate adaptive WAM execution as a future-reality verification problem: the robot should execute longer when the WAM-predicted future remains reliable, and replan earlier when reality deviates from imagination. To this end, we propose Future Forward Dynamics Causal Attention (FFDC), a lightweight verifier that jointly reasons over predicted future actions, predicted visual dynamics, real observations, and language instructions to estimate whether the remaining action rollout can still be trusted.

Geometry Guided Self-Consistency for Physical AI

Dai, Yinwei; Chen, Zhuofu; Yang, Lijie +1 authors

State-of-the-art physical AI models generate a chunk of actions per inference through diffusion or flow matching, iteratively refining an initial noise sample into an action trajectory. We introduce KeyStone, an inference-time self-consistency method for diffusion-based action generation that draws KK candidate action chunks in parallel from a shared model context, clusters them in continuous action space, and returns the medoid of the largest cluster -- no additional model required.

CKT-WAM: Parameter-Efficient Context Knowledge Transfer Between World Action Models

Jiang, Yuhua; Guo, Yijun; Yang, Hongbing +7 authors

We propose \textbf{CKT-WAM}, a parameter-efficient \textbf{C}ontext \textbf{K}nowledge \textbf{T}ransfer framework that transfers teacher WAM's knowledge into a student WAM through a compact context in the text embedding space, rather than output imitation or dense hidden-state matching. Experiments show that CKT-WAM consistently improves zero-shot generalization and achieves the best overall performance on LIBERO-Plus, reaching 86.1\% total success rate with only 1.17\% trainable parameters, while approaching full fine-tuning performance.

EA-WM: Event-Aware Generative World Model with Structured Kinematic-to-Visual Action Fields

Yang, Zhaoyang; Jin, Yurun; Qi, Lizhe +2 authors

To bridge this gap, we present EA-WM, an Event-Aware Generative World Model that effectively closes the loop between kinematic control and visual perception. To fully exploit this geometrically grounded representation, we introduce event-aware bidirectional fusion blocks that modulate cross-branch attention, capturing object state changes and interaction dynamics.

Being-H0.7: A Latent World-Action Model from Egocentric Videos

BeingBeyond Team (Hao Luo; Wanpeng Zhang; Yicheng Feng +6 authors

Being-H0.7 trains a robot policy to anticipate useful future information inside latent queries. A future-aware training branch aligns its hidden states with a deployable branch that sees only current context, avoiding visual rollout during control. Strong benchmark and real-robot results support the complete system, while missing component ablations leave the causal contribution of future alignment unresolved (E02–E09, E12, E15).

Motubrain: An Advanced World Action Model for Robot Control

Motubrain Team

Motubrain combines language, video and action streams in a unified generative robot policy. It transfers video priors into relative end-effector control, then uses asymmetric attention and asynchronous action chunks for deployment. Simulation, world-prediction and real-robot evaluations support different capabilities; the autoregressive attention specification remains internally inconsistent.

Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising

Jun Guo; Qiwei Li; Peiyan Li +7 authors

X-WAM adapts a pretrained video diffusion transformer to predict robot actions, future states and multi-view RGB-D observations together. A one-way depth branch adds geometric supervision, while asynchronous denoising releases actions before completing video generation. The paper reports strong simulated manipulation and reconstruction results plus a small physical earphone-packing evaluation. Its central tradeoff is useful spatial supervision without depth decoding during every action step; limited temporal context and delayed control remain unresolved.

Privileged Foresight Distillation: Zero-Cost Future Correction for World Action Models

Pengcheng Fang; Hongli Chen; Xiaohao Cai

Privileged Foresight Distillation (PFD) compares two attention masks on the same world-action backbone, then learns the future-enabled change in action velocity through an output adapter. Deployment uses only the current frame and corrected action denoising. The reported LIBERO mean improves by 1.15 percentage points over reproduced Fast-WAM, with a measured 5.15% cached-context latency overhead; the title's zero-cost wording does not mean zero additional computation.

RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation

Feng Jiang; Yang Chen; Kyle Xu +8 authors

RoboWM-Bench evaluates whether generated manipulation videos can be converted into robot actions that complete tasks in simulation. Separate human-hand retargeting and robot inverse-dynamics interfaces expose failures hidden by plausible imagery. Reliability tests support these interfaces on real demonstrations, but execution scores still depend on action extraction, reconstructed physics and task-specific checkers.

DexWorldModel: Causal Latent World Modeling towards Automated Learning of Embodied Tasks

Yueci Deng; Guiliang Liu; Kui Jia

DexWorldModel introduces CLWM, which predicts future DINOv3 features and then generates actions conditioned on that prediction. Shared transformer blocks, separate persistent and speculative memories, and asynchronous denoising connect world prediction to robot execution. EmbodiChain supplies synthetic adaptation data. Reported manipulation results are strong, but protocol omissions and a contradictory flow-time convention limit reproducibility.

AIM: Intent-Aware Unified world action Modeling with Spatial Value Maps

Liaoyuan Fan; Zetian Xu; Chen Cao +3 authors

AIM jointly generates future RGB observations, spatial contact maps and continuous robot actions. A masked mixture-of-transformers architecture makes the predicted maps the action head's only route to future information. Supervised learning uses simulator-derived contact labels; post-training freezes world/value prediction and refines actions using map responses plus task rewards. The paper reports 94.0% Easy and 92.1% Hard success on RoboTwin 2.0, but the small post-training improvement, missing mechanism ablations and inconsistent schematic require careful interpretation.

VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis

Xiaolei Lang; Yang Wang; Yukun Zhou +10 authors

VAG synthesizes robot videos and action sequences from an initial image and instruction. A video diffusion transformer supplies pooled clean-latent predictions to an action U-Net during synchronized denoising. The evidence supports improved action prediction and simulation replay, plus a small real-robot policy-pretraining benefit. It does not establish guaranteed alignment or general closed-loop control.

JailWAM: Jailbreaking World Action Models in Robot Control

Hanqing Liu; Songping Wang; Jiahuan Long +9 authors

JailWAM evaluates instruction-induced robot risk by rendering predicted actions as trajectory charts, screening them with a trained vision-language discriminator, and verifying selected candidates in closed-loop simulation. It reports 84.20% human-verified attack success on LingBot-VA, but its cheaper screening pipeline misses some unsafe executions. Success includes motion failure as well as catastrophic risk.

DriveDreamer-Policy: A Geometry-Grounded World-Action Model for Unified Generation and Planning

Yang Zhou; Xiaofeng Wang; Hao Shao +8 authors

DriveDreamer-Policy trains a shared multimodal backbone with depth, video and trajectory generators. Ordered query embeddings transfer geometry and future-scene context into planning without requiring rendered depth or video at inference. Navsim scores and controlled modality ablations support useful joint supervision, while depth evaluation against a learned teacher and incomplete runtime details limit the conclusions.

Enhancing Policy Learning with World-Action Model

Yuci Han; Alper Yilmaz

WAM adds inverse action prediction between consecutive encoder embeddings to a DreamerV2 world model, then trains a separate diffusion policy on its frozen latent features. Table III supports 61.7% versus 45.8% average behavioral-cloning success on eight CALVIN tasks. The mechanism is plausible, but inconsistent headline numbers, baseline names and PPO summaries limit stronger conclusions.

ManipArena: A Controlled Benchmark for Diagnosing Generalization in Real-Robot Manipulation

Yu Sun; Meng Cao; Yang Ping +24 authors

ManipArena evaluates manipulation policies through controlled physical robot trials, task schemas and subgoal scoring. Its strongest lesson is that training recipes and model provenance affect rankings alongside architecture. Language grounding and demonstration selection produce substantial reported gains, but small trial counts, restricted environments and internal reporting inconsistencies limit causal and generalization claims.

Vega: Learning to Drive with Natural Language Instructions

Sicheng Zuo; Yuxuan Li; Wenzhao Zheng +3 authors

Vega learns instruction-conditioned driving by training trajectory denoising together with future-image denoising. Modality-specific transformers exchange information through global causal attention. InstructScene supplies automatically generated instructions describing recorded driving. The strongest reported NAVSIM v2 score uses best-of-six trajectory selection; NAVSIM v1 results are weaker than leading VLA baselines. Future-prediction ablations support visual supervision, while selected images illustrate instruction sensitivity without measuring general instruction-following reliability.

VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs

Haoran Yuan; Weigang Yi; Zhenyu Zhang +9 authors

VTAM adapts a video world model to predict camera and tactile streams, then trains a conditional action expert with an auxiliary deformation-derived force target. It reports large gains in three real robot contact tasks. The strongest evidence is task success and a chip ablation; calibrated force accuracy, broad generalization and the claimed gradient mechanism remain unestablished. Method details below preserve inconsistencies between the formulation and implementation appendix.

Self-Correcting VLA: Online Action Refinement via Sparse World Imagination

Chenyv Liu; Wentao Tan; Lei Zhu +4 authors

SC-VLA adds progress and end-effector-change predictions to a GR00T N1.5-based flow policy, then freezes it and trains a SAC residual controller. Predicted motion supplies a directional reward whose influence decreases with predicted progress. Simulation supports both stages; physical ARX5 trials test only the predictive base policy. The evidence supports improved executed manipulation under the reported protocols, with unresolved reward and evaluation details.

Factored Latent Action World Models

Zizhao Wang; Chang Shi; Jiaheng Hu +4 authors

FLAM learns a video world model whose slots each carry a latent action while sharing interaction-aware dynamics networks. Prediction-trained factorization improves rollouts supplied with future-inferred actions and can supply pseudo action labels for behavior cloning. Its strongest evidence concerns multi-entity video modeling; downstream control is evaluated separately in Procgen.

VideoWorld 2: Learning Transferable Knowledge from Real-world Videos

Zhongwei Ren; Yunchao Wei; Xiao Yu +5 authors

VideoWorld 2 learns discrete visual-dynamics codes with a pretrained diffusion appearance prior, then trains an autoregressive transformer to predict those codes. It improves generated long-horizon craft sequences and transfers latent pretraining to a separately action-supervised CALVIN policy. Its evidence supports benchmark-specific transfer; generated craft success, simulator control and complete appearance disentanglement remain distinct claims.

Wow, wo, val! A Comprehensive Embodied World Model Evaluation Turing Test

Chun-Kai Fan; Xiaowei Chi; Xiaozhu Ju +18 authors

WoW-World-Eval tests whether instruction-conditioned robot videos are visually convincing, task-correct, physically plausible and usable for action extraction. Its 609-sample benchmark combines automated metrics, human judgments and a separate GC-IDM robot-execution test. Hailuo leads the reported aggregate video score, while WoW-wan leads physical execution. The central lesson is that a plausible imagined manipulation and an executable one are distinct outcomes; calibration and evaluator dependence constrain how broadly the scores can be interpreted.

Dream2Flow: Bridging Video Generation and Open-World Manipulation with 3D Object Flow

Karthik Dharmarajan; Wenlong Huang; Jiajun Wu +2 authors

Dream2Flow converts generated human-interaction videos into 3D object trajectories, then uses domain-specific optimization or reinforcement learning to make a robot realize them. Its central benefit is an object-level interface across embodiments; its reliability still depends on video geometry, tracking, contact assumptions and the downstream controller.

CoVAR: Co-generation of Video and Action for Robotic Manipulation via Multi-Modal Diffusion

Liudi Yang; Yang Bai; George Eskandar +5 authors

CoVAR co-generates instruction-conditioned video and robot actions through separate, interacting diffusion transformers. Bridge Attention connects a pretrained video branch to an action branch; a UNet decodes actions, with an additional refinement model for LIBERO90. Reported simulation and UR5 successes support the complete system, while refinement dependence, incomplete evaluation details, and roughly four-second sequence generation limit broader conclusions.

WorldLens: Full-Spectrum Evaluations of Driving World Models in Real World

Ao Liang; Lingdong Kong; Tianyi Yan +19 authors

WorldLens evaluates driving video generators through appearance, reconstructability, planner behavior, perception and human judgment. Its strongest lesson is that favorable image metrics coexist with poor closed-loop route completion. A separate LoRA-trained critic learns score-and-rationale outputs from human annotations; its generalization evidence remains qualitative.

VideoVLA: Video Generators Can Be Generalizable Robot Manipulators

Yichao Shen; Fangyun Wei; Zhiying Du +5 authors

VideoVLA adapts CogVideoX-5B into a robot policy that jointly denoises future video latents and executable action chunks, conditioned on language and the current image. Its clearest gains concern novel objects and skills transferred between embodiments. Video supervision matters strongly in the reported ablations, but imagined success exceeds executed success and deployment remains slow.

MindDrive: An All-in-One Framework Bridging World Models and Vision-Language Model for End-to-End Autonomous Driving

Bin Sun; Yaoguang Cao; Yan Wang +8 authors

MindDrive couples action-conditioned BEV scene prediction with anchor refinement and a VLM trajectory scorer. Its main contribution is the connection between future-aware candidate generation and selection. NAVSIM scores support planning gains, while incomplete training specifications and inconsistent result reporting limit reproducibility.

Video2Act: A Dual-System Video Diffusion Policy with Robotic Spatio-Motional Modeling

Yueru Jia; Jiaming Liu; Shengbang Liu +7 authors

Video2Act turns observed video into conditioning for a separate robot action policy. Sobel filtering emphasizes spatial structure in video-model features, while temporal Fourier filtering emphasizes motion. A slow video network refreshes these representations for a faster diffusion action head. The strongest evidence is improved executed manipulation on RoboTwin and a small real-robot evaluation; the speed claims require separating action-chunk throughput from fresh-feedback control.

LatBot: Distilling Universal Latent Actions for Vision-Language-Action Models

Zuolei Li; Xingyu Gao; Xiaofan Wang +1 authors

LatBot learns scene and motion latents from language-conditioned robot and human manipulation videos, supervises them through future-image and physical-action decoding, then distills them into a VLA student. An action expert converts the student's features into executable commands. The strongest evidence concerns downstream manipulation success; universal physical understanding and preserved reasoning remain broader interpretations of these results (e03–e06, e12–e16).

Learning Massively Multitask World Models for Continuous Control

Nicklas Hansen; Hao Su; Xiaolong Wang

Newt extends TD-MPC2 into a language-conditioned agent that learns latent dynamics from demonstrations and online interaction across MMBench. Its practical recipe pretrains the world model and policy prior, retains action supervision, and uses the model to plan. It improves over the evaluated multitask baselines, but specialist policies remain stronger and long open-loop execution is uneven. The evidence concerns simulated, predominantly state-based control.

RynnVLA-002: A Unified Vision-Language-Action and World Model

Jun Cen; Siteng Huang; Yuqian Yuan +11 authors

RynnVLA-002 finetunes a shared Chameleon backbone for action prediction and action-conditioned image prediction, then adds a parallel continuous-action head. Its strongest evidence is mutual training benefit: better executed policies and better held-out visual predictions. Policy inference uses no imagined-image rollout. The reported 97.4% LIBERO average is competitive, while physical evidence is limited to SO100 pick-and-place.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

John Won; Kyungmin Lee; Huiwon Jang +2 authors

DUST augments a frozen vision-language backbone with jointly denoised actions and future visual embeddings. Separate streams exchange information through attention; independent noise levels and unequal sampling budgets accommodate their different dynamics. Controlled comparisons support improved manipulation, but do not establish general causal world understanding.

Generative World Modelling for Humanoids: 1X World Model Challenge Technical Report

Riccardo Mereu; Aidan Scannell; Yuxin Hou +6 authors

Team Revontuli uses two separate, state-conditioned predictors: a LoRA-adapted Wan video model for future RGB frames and a spatio-temporal Transformer for future discrete tokens. Both lead the reported challenge leaderboard, but their scores measure conditional prediction rather than executed humanoid control.

MimicDreamer: Aligning Human and Robot Demonstrations for Scalable VLA Training

Haoyun Li; Ivan Zhang; Runqi Ouyang +12 authors

MimicDreamer converts human demonstrations into robot-looking videos paired with retargeted actions, then post-trains a separate π0 policy. The contribution is training-data alignment across viewpoint, embodiment and appearance. Physical-task results improve with added synthetic demonstrations, but real-robot supervision remains part of the evaluated pipeline and several headline summaries conflict with the tables.

Vidar: Embodied Video Diffusion Model for Generalist Manipulation

Yao Feng; Hengkai Tan; Xinyi Mao +5 authors

Vidar adapts an embodied video generator to a target bimanual robot, then converts predicted imagery into controls using a separately trained masked inverse dynamics model. Low-demonstration real-world results and component ablations support this design; action observability, open-loop execution, and substantial prior training constrain its generality.

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge

Wenyao Zhang; Hongsi Liu; Zekun Qi +11 authors

DreamVLA trains a shared backbone to anticipate motion-relevant regions, depth and semantic features, then conditions action diffusion on its latent predictions. Explicit visual decoders disappear at inference. It reports 4.44 completed instructions on CALVIN ABC-D and 76.7% real-robot success under an attempt-limited protocol. Its lesson is selective forecasting, tempered by unresolved mask, configuration and table inconsistencies.

RoboScape: Physics-informed Embodied World Model

Yu Shang; Xin Zhang; Yinzhou Tang +4 authors

RoboScape predicts action-conditioned robotic RGB and depth videos using coupled autoregressive branches. Depth-feature feedback and motion-selected keypoint supervision aim to improve geometry and interaction dynamics. Reported benefits include stronger video metrics, useful synthetic policy-training data, and correlated simulator-based policy evaluation; the ablations expose tradeoffs, and physical robot deployment remains future work.

AMPLIFY: Actionless Motion Priors for Robot Learning from Videos

Jeremy A. Collins; Loránd Cheng; Kunal Aneja +3 authors

AMPLIFY learns a compact vocabulary of visual motion from point tracks, predicts that motion from an image and task instruction, and translates it into robot actions through a separate inverse model. Its strongest evidence concerns scarce target-task action labels, including transfer where target-task videos remain available. Better track prediction and physical task completion are evaluated separately.

3DFlowAction: Learning Cross-Embodiment Manipulation from 3D Flow World Model

Hongyan Zhi; Peihao Chen; Siyuan Zhou +4 authors

3DFlowAction learns instruction-conditioned object trajectories from human and robot videos, checks a rendered endpoint with GPT-4o, and converts accepted flow into robot poses through grasp selection and optimization. Its four-task physical evaluation reports 70% success, but transfer depends on rigid grasp geometry and follows task-specific human-video fine-tuning.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction

Chenyou Fan; Fangzheng Yan; Chenjia Bai +4 authors

CogRobot adapts video diffusion to bimanual control through three learned components: instruction-conditioned flow prediction, flow-conditioned RGB prediction, and a task-specific goal-reaching action policy. Flow offers an intermediate description of motion without requiring action labels for the video models. The strongest physical result is Pull Box success of 0.75 versus 0.05 for DP, but the evidence does not establish a universal action policy or broad unseen-task transfer (e02-framing, e05-video, e06-controller, e09-real, e13-limit).

RoboTransfer: Controllable Geometry-Consistent Video Diffusion for Manipulation Policy Transfer

Liu Liu; Xiaofeng Wang; Guosheng Zhao +7 authors

RoboTransfer augments robot demonstrations by changing their visual appearance while conditioning video diffusion on the demonstrated geometry. Jointly encoded camera views, metric depth, normals and separate background/object references support multi-view synthesis. A separately trained ACT policy benefits from the augmented observations on two physical manipulation tasks. This is evidence for offline data augmentation, with remaining uncertainty about statistical reliability and physical fidelity.

Learning Generalizable Robot Policy with Human Demonstration Video as a Prompt

Xiang Zhu; Yichen Liu; Hezhong Li +1 authors

A human demonstration video conditions a frozen video model whose features guide a separate diffusion policy for an Xhand robot. Human hand motions are reconstructed and retargeted into the policy's action space, allowing human demonstrations to supervise action learning. The reported transfer concerns objects and skills absent from robot training demonstrations but present in human training data; it does not establish learning an entirely untrained skill from the inference prompt alone.

WorldEval: World Model as Real-World Robot Policies Evaluator

Yaxuan Li; Yichen Zhu; Junjie Wen +2 authors

WorldEval estimates robot-policy rankings by turning internal policy embeddings into generated manipulation videos and judging their outcomes. Policy2Vec conditions a separately adapted WAN 2.1 simulator; Gemini-2.0 supplies success labels. Paired experiments support relative-ranking usefulness on the tested tabletop setup, while imperfect action fidelity, checkpoint-distribution dependence and inconsistent source reporting limit stronger conclusions.

FLARE: Robot Learning with Implicit World Modeling

Ruijie Zheng; Jing Wang; Scott Reed +18 authors

FLARE trains a flow-matching robot policy to match future observation embeddings inside its action-denoising transformer. Compact, action-trained visual-language targets supply an auxiliary learning signal, including from action-free human videos. Simulation and real-robot improvements support this training recipe; they do not establish an explicit planner or calibrated world simulator.

EnerVerse-AC: Envisioning Embodied Environments with Action Condition

Yuxin Jiang; Shengcong Chen; Siyuan Huang +8 authors

EVAC turns robot action sequences into future camera observations using a video diffusion model. Spatial pose maps, temporal action differences and camera rays condition the generated environment. A separate policy can interact with that environment for evaluation, or learn from synthetic trajectories. The strongest evidence is a small policy-data augmentation experiment and agreement with real-robot evaluation trends; physical accuracy remains incompletely measured.

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models

Yue Hu; Siyuan Huang; Yue Liao +5 authors

EWMBench evaluates instruction-conditioned robot videos through scene consistency, end-effector motion, and semantics. Its seven-model comparison favors domain-adapted generators, while controlled trajectory corruptions expose why static-looking plausibility is insufficient. These are offline video-evaluation results, with unresolved reporting inconsistencies, rather than demonstrations of executed robot control.

UniVLA: Learning to Act Anywhere with Task-centric Latent Actions

Qingwen Bu; Yanting Yang; Jisong Cai +5 authors

UniVLA learns a discrete action vocabulary from language-annotated videos, teaches a vision-language policy to predict that vocabulary, and adapts a visual-conditioned head to physical controls. Its distinguishing idea is to separate task-relevant motion from distracting changes before policy pretraining. Strong manipulation and navigation results support transfer, but deployment still needs action-labeled adaptation; future-observation prediction trains the latent representation rather than serving as the evaluated inference-time planner (e-lam, e-policy, e-decode, e-libero, e-nav, e-limitations).

CLAM: Continuous Latent Action Models for Robot Learning from Unlabeled Demonstrations

Anthony Liang; Pavel Czempin; Matthew M. Hong +4 authors

CLAM learns continuous action codes from robot observation transitions, grounds them with limited action-labeled data, then imitates expert videos in that latent space. Deployment uses a policy and action decoder; future-observation prediction serves training. Its strongest evidence concerns action-label scarcity within one robot embodiment, with several unresolved protocol details.

CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models

Qingqing Zhao; Yao Lu; Moo Jin Kim +12 authors

CoT-VLA makes a 7B multimodal model generate a future subgoal image before predicting a robot action chunk. The shared model learns image prediction from demonstrations and captioned videos, while action supervision requires demonstrations. Experiments support useful goal conditioning and stronger average manipulation performance, with uneven task gains, substantial inference overhead and reporting inconsistencies. Visual chain of thought here means an explicit predicted image used for control; general reasoning is not independently established.

LUMOS: Language-Conditioned Imitation Learning with World Models

Iman Nematollahi; Branton DeMoss; Akshay L Chandra +3 authors

LUMOS trains a language-guided manipulation policy by practicing inside a world model learned from offline robot play. Frozen recurrent dynamics support actor-critic learning with expert-latent matching rewards; hindsight plans and language alignment organize behavior. CALVIN and physical tabletop experiments support this combination, with modest gains over adapted HULC and larger component-ablation losses. Transfer means deployment after learning from that environment’s offline data, without online policy fine-tuning.

VILP: Imitation Learning with Latent Video Planning

Zhengtong Xu, Qiang Qiu, Yu She

VILP learns observation-conditioned future videos in a compressed latent space, decodes them, and maps adjacent frames to actions through a separate low-level policy. This makes short-horizon video planning practical on the tested tasks and lets task videos supply information beyond scarce action labels. Simulation gains depend on data and evaluation protocol; real-robot evidence comprises 15 trials per method.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos

Zhongwei Ren; Yunchao Wei; Xun Guo +4 authors

VideoWorld predicts compact multi-step dynamics codes alongside video frames, then converts predictions into task operations. On rendered 9×9 Go and simulated robot tasks, this representation substantially improves over video-only prediction. Expert-curated training data, language conditioning and a separately supervised robot action decoder qualify the broader claim of learning solely from unlabeled videos.

EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation

Siyuan Huang; Liliang Chen; Pengfei Zhou +8 authors

EnerVerse learns instruction-conditioned, multi-view future-video representations, then conditions a separate action diffusion head on their backbone features. Sparse history supports long tasks; an offline 4D Gaussian Splatting loop refines training videos. Its strongest benchmark average uses depth-rendered auxiliary views, while physical block placement exposes an instruction-following weakness.

Scaling Offline Model-Based RL via Jointly-Optimized World-Action Model Pretraining

Jie Cheng; Ruixi Qiao; Yingwei Ma +5 authors

JOWA learns an Atari world model and distributional Q-function through one shared transformer, then searches short imagined futures to choose actions. Its strongest evidence is improved aggregate game return and data-efficient offline adaptation; neither universal game-wise scaling nor unconditional planning optimality is established (e02, e03, e04, e07, e08, e11, e12).

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Homanga Bharadhwaj; Debidatta Dwibedi; Abhinav Gupta +7 authors

Gen2Act turns a language instruction and an initial scene image into a generated human demonstration, then uses that video to condition a separate closed-loop robot policy. A pretrained VideoPoet supplies the demonstration without robot-specific fine-tuning. Auxiliary point-track prediction teaches policy representations to retain motion cues, while deployment predicts actions directly from video features and recent robot observations. Real robot results support improved generalization relative to the reported baselines, but plausible generation does not ensure correct execution, and long-horizon reliability remains limited.

DynaMo: In-Domain Dynamics Pretraining for Visuo-Motor Control

Zichen Jeff Cui; Hengkai Pan; Aadhithya Iyer +2 authors

DynaMo uses action-free dynamics pretraining to make visual features useful for imitation learning. An encoder, latent inverse model and forward model jointly learn to predict next-frame embeddings; a separate policy then learns from frozen features and labeled demonstrations. Results favor DynaMo on several manipulation tasks, but include ties and initialization regressions. The contribution is a representation-learning objective, with no demonstrated use of the dynamics models for online planning.

This&That: Language-Gesture Controlled Video Generation for Robot Planning

Boyang Wang; Nikhil Sridhar; Chao Feng +4 authors

This&That turns an initial image, language and pointing coordinates into a generated video plan, then uses a separate DiVA behavior-cloning controller to follow it with live feedback. Gestures improve intent disambiguation in Bridge video generation and simulated block manipulation; real-robot execution remains untested.

Dreamitate: Real-World Visuomotor Policy Learning via Video Generation

Junbang Liang; Ruoshi Liu; Ege Ozguroglu +5 authors

Dreamitate learns task-specific video generation from human tool demonstrations, then extracts tool poses from synthesized stereo videos for robot execution. It outperforms the tested Diffusion Policy baseline on four physical manipulation tasks, while depending on calibrated cameras, known rigid tools, and slow open-loop generation. The central evidence concerns executed tool trajectories under specified object and scene shifts, rather than unrestricted manipulation or a general action-conditioned simulator.

RoboDreamer: Learning Compositional World Models for Robot Imagination

Siyuan Zhou; Yilun Du; Jiaben Chen +3 authors

RoboDreamer composes phrase-conditioned video diffusion predictions to imagine robot plans for unfamiliar instruction combinations. Optional goal images or sketches sharpen spatial specifications; a separate inverse-dynamics model converts imagined frames into actions. The strongest language-only evidence concerns human-rated video alignment, while executed success is measured separately in RLBench simulation (E03, E10–E13).

Learning Universal Policies via Text-Guided Video Generation

Yilun Du; Mengjiao Yang; Bo Dai +5 authors

UniPi turns a language instruction and current image into a video plan, refines its timing, then uses a separately trained inverse model to produce robot controls. Simulated manipulation results support compositional and multitask transfer. Internet pretraining improves generated real-scene plans, but its reported success is a classifier judgment on imagined final frames, not measured physical execution. [e02, e04, e06, e10, e14, e16]

Missing a paper, or found a correction?

Suggest an update on GitHub ↗