Reading reports

These reading notes explain each paper’s mechanism, connect results to their source evidence, and record the limits of what was reviewed.

About these reading notes
  • The idea: a concrete problem and a step-by-step explanation.
  • The evidence: results with their evaluation setting and source location.
  • The boundaries: reading coverage, limitations and reproduction questions.

AI-assisted research notes. Verify claims against the linked primary sources.

Catalog entries
564
Illustrated reports
558
Full-text reviewed
553
Partial / resource reviews
11

Report directory

Updated 11 Sept 2026

558 illustrated readings and 6 text reports are available. The status below records source reading scope. Illustrated editions are labeled separately; remaining entries keep their place in the queue.

564 entries

Source access is shown separately from reading status.

Programmable World Model

Zheng-Hui Huang; Guixu Lin; Jiacheng Lin; Yi-Chuan Huang; Ruihan Yu; Muyao Niu; Siqi Yang; Yu-Lun Liu; Yung-Yu Chuang; Kaipeng Zhang; Zhixiang Wang

Programmable World Model makes a rule-executing engine authoritative for world facts and uses a conditioned video model to render them. State-augmented 3D oriented bounding boxes connect these components without requiring detailed animated assets. On CombatStateBench, the system reports 94% visible-alive-count accuracy and 98% death-state accuracy, but these permissive global metrics do not establish correct entity-specific interactions or autonomous action selection (e-world, e-controls, e-table, e-metrics).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

DUET-DINO: Simultaneous Cross-View World Modeling for Latent Planning in Robot Manipulation

Nisarga Nilavadi; Ralf Römer; Moritz Reuss; Michael Krawez; Tobias Jülg; Angela P. Schoellig; Rudolf Lioutikov; Wolfram Burgard

DUET-DINO couples side- and wrist-camera latent predictors through cross-attention, then searches seven-dimensional robot actions against paired goal images. Its clearest evidence is improved simulated spatial and orientation planning; hardware success remains limited under a smaller planning budget and safety terminations (e-architecture, e-planning, e-angled, e-hardware).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

FolDeX: A Physical-World Benchmark for Long-Horizon Robotic Manipulation of Deformable Objects

Chenhuan Liu; Yi Xu; Feng Wu; Hanyang Wang; Wenxiao Kuai; Weihao Ding; Shan Wang; Yang Liu; Shuyong Gao; Wenqiang Zhang

FolDeX evaluates complete physical garment-folding episodes and asks whether expensive robot experience can be reused across recovery states, tasks, scenes, and embodiments. Its reference policies use π0. Recovery-augmented training raises reported average success from 80.75% to 95.00%, but transfer studies remain preliminary and the missing quality-scoring appendix prevents independent reconstruction of FoldScore from the supplied paper alone.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Grounding Generated Video Plans in Simulation Towards Versatile Dexterous Controllers

Tianyue Wu; Boyuan An; Shuqi Zhao; Heyu Guo; Wanli Xing; Yi Ma; Kaifeng Zhang; Ruihai Wu; Masayoshi Tomizuka

GALATEA turns image-and-language-conditioned manipulation videos into 3-D hand–object references, then learns simulation-based controllers that physically track them. Its central interface retains finger motion and object motion together. Multi-skill experts are distilled into one feedback policy; simulated transfer and 27/40 real-world successes on unseen plans support useful, but limited, generalization (e02, e09, e14, e17).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

HaWMPO: Hallucination-Aware World Model-based Policy Optimization for Generalist Robot Policy

Zengjue Chen; Peidong Liu; Jiawei Li; Qi Wang

HaWMPO post-trains OpenVLA-OFT inside a frozen video world model. A separate hallucination detector discounts rewards assigned to unreliable imagined action chunks before GRPO updates the policy. Table 1 reports 63.7% average LIBERO success, but gains vary by suite and the penalty-selection description is inconsistent. Physical testing provides preliminary evidence on two G1 tasks.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

JEPA Policy: Diffusion-Free Imitation Learning via Paired Action and Future Representation Prediction

Jie Xu; Kangjin Yu; Ziyi Jin; Junjie Gao; Liqing Chen; Yixian Li; Shuai Tian; Zhongpu Xia

JEPA Policy trains a shared Transformer to predict an action chunk and the visual representation observed later in the same demonstration. Two deterministic passes refine both outputs. Its strongest evidence is improved simulated control relative to action-only MIP, supported by topology controls; future error also offers a delayed, task-dependent rollout diagnostic.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Compact Visuotactile World Models for Lifting: Prediction, Reward Alignment, and Force Constraints

Qinzhen Ma; Sida Peng

This study asks whether better contact prediction produces better lifting decisions. A compact visuotactile model improves forecasts over vision alone, but persistence challenges the value of its dynamics. An exploratory reward revision improves executed simulator lifting, while simple force feedback remains stronger under a force budget. A separate GelSight experiment exposes the gap between frame and trajectory reliability.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Valerant: An Automatic Navigable Game Map Generator via Action-Conditioned World Model Exploration

Yiran Qiao; Feng Wang; Jing Ma

Valerant expands one game screenshot into a persistent 3D point-cloud map by generating alternative action-conditioned videos, reconstructing each with SLAM, and committing the best admissible branch. A frozen video model supplies futures; an external exploration policy selects actions. Perceptual comparison and qualitative ablations support this prototype, while geometric reliability and reproducibility remain incompletely measured.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

SyncWorld: Visual Calibration Enables World Models as Zero-Shot Simulators

Yuncong Yang; Zhengtao Han; Furkan Ozyurt; Zeyuan Yang; Han Yang; Junyi Cao; Haoyu Zhen; Yilun Du; Chuang Gan

SyncWorld interprets robot commands through a short visual calibration context, then predicts proposed actions’ consequences with a video diffusion model. Calibration-conditioned training and distillation also support history-only prediction. It improves held-out video metrics and selected LIBERO policy outcomes, but the evidence is strongest for short-horizon simulation: ranking quality, object interactions and unseen-embodiment controllability remain limiting factors.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

WorldAgen: Unified State-Action Prediction with Test-Time World Model Training

Chi Wan; Kangrui Wang; Yuan Si; Pingyue Zhang; Manling Li

WorldAgen shares a Transformer between action prediction and future-observation prediction, then adapts the shared representation using exploratory target-environment transitions. Simulated manipulation improves after observation-only test-time training (TTT), while implementation ambiguities and unreported uncertainty limit the strength of the generalization claim.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model Pretraining

Yuran Wang; Siqiao Huang; Mingleyang Li; Chenhao Zhang; Jiaqi Liang; Weiyang Jin; Yue Chen; Xuemin Chi; Donghao Zhou; Qize Yu; Yu-Kai Wang; Yuhan Rui; Shenzhe Yao; Zhen Yuan; Zhenhao Shen; Kefei Zhu; Zijie Zhu; Ning Gao; Xiaowei Chi; Guanqi He; Shanghang Zhang; Hao Dong; Lin Shao; Hang Zhao

OpenWAM turns world–action modeling into a modular design study, then instantiates OpenWAM-α: a video DiT coupled to a dedicated ActionDiT through mutual attention and joint denoising. Its strongest lesson is conditional: embodied pretraining improves scene transfer, but strong manipulation scores coexist with substantial visual-robustness failures. The evidence below separates controlled ablations, final-model benchmarks and physical execution.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Beyond Task Success: Stage-Wise Reliability of World Model Planning under Sensing Degradation

Geonmyeong Lee; Byoung-Tak Zhang

This diagnostic study follows synthetic sensing disturbances through an existing world-model planner's representation, action-conditioned prediction, candidate ranking, and executed task outcome. Paired OGBench evaluations show that a disturbance's relative impact can change between stages. Temporal information position matters even when aggregate representation shifts are similar. These diagnostics locate sensitivity; they neither establish a general predictor of task failure nor demonstrate successful sensing mitigation.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

WM-Craftnet: World Synesthesia Model for Generalizable and Robust Dexterous In-Hand Manipulation

Jie Yin; Zeyuan Zhao; Xiaojing Tan; Yang Liu; Chiyu Wang; Xinyang Gu

WM-Craftnet learns a predictive visuotactile state that conditions a separate dexterous manipulation policy. Its main contribution is recurrent perception for executed control: noisy depth, touch, proprioception, and action history inform a latent state trained with clean-depth and other prediction targets. Hardware rotation improves substantially, but severe-disturbance recovery remains limited (e-rssm, e-policy, e-stress).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Learning to Use Imagination: Progress-Conditioned Future Utilization for World Action Models

Yijie Zhu; Zitong Yu; Wei Li; Hui Ma; Wen Li; Rui Shao; Liqiang Nie

ProWAM learns execution progress from recent action–observation feedback and recurrent memory, then uses that state to redistribute imagined-future attention inside a joint video–action policy. Its strongest evidence combines executed manipulation results with ablations separating progress estimation from future utilization.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

How to Learn from What a Human Would Avoid? Intervention-Aware World Models with Real-World RL for Dexterous Manipulation

Jiaju Yin; Zhenhui Zhang; Lixin Xu; Heng Zhang; Jun Shao; Yating Feng; Arash Ajoudani; Renjing Xu

WHIRL learns when a human would take over a dexterous robot and uses that prediction to steer residual-policy training. A frozen imitation prior supplies nominal behavior; an action-conditioned, one-step latent world model supports critic learning and actor risk shaping. Five real-robot tasks show higher autonomous success, but the evidence is restricted to one training seed, one operator and familiar workspace regions.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Learning Counterfactual World Models for Embodied Reasoning under Partial Observability

Todd Y. Zhou; Daniel Zhang

Counterfactual Latent World Models (CLWM) train action-conditioned imagined futures to preserve differences in intervention outcomes even when observations look alike. A recurrent world model supplies latent rollouts to model-predictive control; privileged outcome labels supervise an additional contrastive objective during training. Reported simulation gains reach 11.6 percentage points in navigation success. The evidence supports targeted representation training, while supervision-matched controls and audits of pretrained encoders remain open.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation

AgiBot Research Team

GE-Act 2.0 connects a compact visual representation, a one-step future generator, and a separate inverse dynamics model. Its central training problem is pairing generated futures with compatible recorded actions. KASO selects futures using the current action model before updating both modules. Real-robot experiments show broader manipulation capability as co-training data grows, with persistent weaknesses in fine manipulation and ordinal language; simulation comparisons use a separate adaptation protocol.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

TacPAC: Tactile Prediction and Real-Time Action Correction in World-Action Models for Contact-Rich Manipulation

Zipei Ma; Xiaofei Wei; Junzhe Jiang; Shunlin Lu; Li Zhang

TacPAC predicts future visual and tactile observations while planning an action chunk, then uses incoming touch to revise its unfinished actions against a fixed prediction-and-plan cache. On five physical manipulation tasks, it reports 64% average success versus 48% for T-Rex and 22% for its vision-only ablation. The central evidence is the controlled cache-access ablation, not visual prediction quality alone.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

TourPhysics: Bringing Physics to World Models for Exploration and Manipulation from a Single Image

Xin Zhang; Yabo Chen; Zixuan Duan; Haibin Huang; Chi Zhang; Feng Xu; Xuelong Li

TourPhysics renders persistent camera tours and object manipulations from a single image plus a declared physical scene. A simulator fixes each trajectory before diffusion generates its appearance; accepted windows publish physical endpoints and visual memory together. The strongest evidence concerns prescribed motion and modest improvements on held-out revisits, with limited physical-event support and no real-world validation.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

GIFT: Guided Intermediate Feature Training via Action-Oriented Structural Supervision for Robotic Manipulation

Yupeng Zheng; Xiang Li; Songen Gu; Yuhang Zheng; Shuai Tian; Weize Li; Linbo Wang; Chaoyue Li; Qichao Zhang; Haoran Li; Zhongpu Xia; Ya-Qin Zhang; Shuicheng Yan; Dongbin Zhao

GIFT trains robot-policy features to retain geometry, object–end-effector relations and instruction-relevant regions. The same auxiliary objectives improve three different action formulations while their default deployment uses no auxiliary predictions. Evidence is strongest for matched-baseline robustness gains, with substantial annotation requirements and uneven benefits across shifts.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Toward Unified Robot Learning: Bridging Representation, Vision-Language-Action, and World Models

Shaunak A. Mehta; Ananya Hazarika; Haochen Zhang; Fan Yang; Ryo Moriyama; Wenkai Li; Yash Patel; Kanata Suzuki

This survey organizes robot learning around representations that encode the environment, VLA policies that generate actions, and world models that predict consequences. Its useful contribution is a vocabulary for tracing how information and feedback cross those interfaces. Integration can remain modular; the authors argue that its value should be demonstrated through improved behavior under uncertainty, distribution shift and temporal dependencies. A small original CALVIN video-prediction diagnostic illustrates hidden-object and collision failures, but supplies no quantitative proof that a unified architecture resolves them.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

WISE: World-model-guided Imagination Scheduling for Efficient Post-training of Vision-Language-Action Models

Chenhao Zhang; Hanyu Zhao; Hang Cheng; Tengfei Pan; Long Zeng

WISE post-trains a VLA action head by selectively imagining alternative behaviors at visually identified interaction states. A separate frozen world model predicts bounded futures; a frozen evaluator ranks them; updates supervise only the first action chunk from a real context. The strongest controlled evidence is the scheduling ablation: higher task success with substantially less imagination computation. The method, results, and reproduction boundaries below trace this conclusion to the primary text.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

SV-WAM: An Efficient Surround-View World-Action Model for End-to-End Autonomous Driving

Jinyang Wang; Shiwei Li; Junjian Wang; Zhiqiang Deng; Jianbin Gao; Yihang Zhao; Liu Liu; Yongjia Zhao; Jinlong Chen; Huirui Xu; Yifeng Pan; Kangwei Liu; Fan Ren; Ji Tao; Minghao Yang

SV-WAM trains one shared transformer to denoise driving actions and future surround-view video, but blocks action attention to future-video tokens. Deployment retains six-camera history while generating only trajectories. A differentiable vehicle-footprint loss improves road compliance. The strongest reported result is 91.0 EPDMS on NAVSIMv2 navtest; the evidence concerns benchmark planning, with weaker hard-split performance and no demonstrated real-vehicle deployment.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Drive-HWM: Hierarchical World Models for Dynamic-Latent Guided Autonomous Driving

Zhaoxin Fan; Tianbao Zhang; Wenjun Wu; Xiaofeng Wang; Yeying Jin; Jian Zhao; Zheng Zhu; Shuicheng Yan

Drive-HWM couples a periodically refreshed predictor of future motion latents with an observation-grounded autoregressive driving policy. Optical flow supervises the slow branch; next-frame RGB tokens supervise the fast branch during training. NAVSIM tables report stronger aggregate driving scores, but conflicting prose and incomplete implementation details limit precise reproduction.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Toward Physically Grounded JEPA World Models for Goal-Conditioned Robotic Planning

Muyuan Liu; Yue Huang; Zheng Liang; Xiang Gao

SA+IDM trains an action-conditioned JEPA world model with two auxiliary heads: inverse dynamics recovers executed actions, while state alignment predicts measured physical state from consecutive image representations. Deployment uses only the encoder and latent predictor inside CEM planning. State alignment improves all four reported tasks over IDM alone, while the diagnostics challenge average temporal straightening as a sufficient representation-quality criterion (e2–e12).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Building Pretraining Data for World Models: An Unreal Engine-Based Pipeline for Action-Conditioned Video Generation

Haoyu Wang; Songchun Zhang; Haoran Li; Haoyang Huang; Zeyue Xue; Nan Duan

This paper documents the Unreal Engine synthetic-data component used in EchoWM: simulate character motion once, record controls and states, then replay the trajectory for high-quality multi-view rendering. Its contribution is a production system combining scene curation, cache locality and failure recovery. Reported video volume establishes production scale; downstream learning utility remains untested here. [e-scope, e-workflow, e-scale, e-limits]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Latent Energy Action Planning with World Models

Phu Pham and Aniket Bera

LEAP refines an action horizon through frozen LeWorldModel dynamics, combining latent-goal matching with a learned decoder’s terminal-state error. A trained proposal initializes search; projection bounds executed controls. Four-domain mean success rises from 77.5% to 94.8% against matched LeWM+CEM. The narrower energy ablation improves from 91.0% to 96.5%, separating the extra objective’s contribution from the complete planner change. Numerical goal descriptors and trustworthy learned rollouts remain prerequisites (e-energy, e-proposal, e-optimization, e-main, e-ablation, e-limits).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Spatially Aware World Action Model via Geometric Latent Diffusion

Gonzalez, Javier Alejandro Lopetegui; Pacaud, Paul; Schmid, Cordelia

SA-WAM adds metric depth to a pretrained video diffusion policy by encoding log-normalized depth through the same frozen tokenizer used for RGB. A shared transformer jointly predicts actions and future RGB-D observations. Its strongest evidence combines a controlled depth-normalization ablation, improved simulated task success, and physical UR5 evaluations scored with partial credit. The prediction-error analysis establishes an association with failures, without demonstrating a deployed failure detector. [e02, e04, e09, e10, e14, e18]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

World-Coherent Decoding: Self-Verifying Test-Time Planning for World Action Models

Zhang, Chuhan; Ito, Seiji; Hoshino, Kenta; Ikehata, Satoshi; Sato, Ikuro

World-Coherent Decoding selects among a frozen WAM's imagined futures and associated action chunks, then uses observed prediction errors to improve later selections. It raises simulated Hard success from 55.80% to 60.90% under limited randomized-scene supervision. Its central contribution is a feedback-calibrated decoding procedure; physical-robot evidence is qualitative, and missing implementation appendices constrain reproducibility. [E03, E08, E11, E14, E19]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Latent Action as Intention Enables Efficient Future Imagination for World Action Models

Li, Xiang; Zheng, Yupeng; Gu, Songen; Ma, Huailiang; Yu, Feng; Zheng, Yuhang; Nie, Xian; Yuan, Shanshuai; Zang, Yujie; Li, Weize; Tian, Shuai; Liu, Moyang; Zhang, Ya-Qin; Ding, Wenchao

LAWA retains test-time future prediction through a compact sequence of transition embeddings, jointly denoised with robot actions. A frozen tokenizer supplies discrete training targets, but deployment uses continuous latent states without quantization. Future-video prediction remains a training objective. The complete system improves manipulation success over matched Fast-WAM while reducing latency relative to Joint-WAM; egocentric pre-training is central to this trade-off. [E6, E7, E11, E16, E17]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training

Wu, Xionghao; Yang, Yijun; Zhou, Shiyang; Sun, Haoze; Liu, Jianhui; Yu, Songsong; Zhang, Jiyao; Li, Wenbo; Wang, Bo; Ma, Guoqing; Song, Lin; Liao, Renjie; Zheng, Shenghe; Tang, Wei; Qi, Xiaojuan; Li, Yanwei; Zhang, Yuan; Tian, Zhuotao; Huang, Haoyang; Duan, Nan

ZimaBlue converts action-free embodied video into robot control through video pre-training, cross-embodiment video–action alignment, and target-domain adaptation. Its deployed controller combines a predictive Slow transformer with a smaller Fast action generator that consumes cached video features and fresh observations. The central evidence is executed Franka manipulation: expanding training raises reported task-macro success from 36.1% to 77.8%. A separately accelerated configuration reaches 33.0 ms inference latency with 75.0% success. The work supports video scaling and asynchronous feedback as useful ingredients, while leaving the specific necessity of predictive cached features insufficiently isolated. [e03, e05, e07, e10, e12, e13]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Motus2: A Self-Evolving General World Model for Dexterous Manipulation

Bi, Hongzhe; Zhou, Zihao; Tang, Yihang; Pang, Jingrui; Huang, Shuhe; Liu, Haitian; Wang, Runqing; Huang, Shuai; Wang, Yichen; Cheng, Yiming; Zhao, Ruowen; Li, Zhenghua; Tan, Hengkai; Liu, Xiaolong; Wan, Jinhui; Liu, Jiabao; Zhao, Min; Bao, Fan; Zhu, Jun

Motus2 adapts a jointly pretrained video–action transformer into an action-first controller: propose an action chunk, optionally simulate its visual consequences, and evaluate the branch using learned relative progress. Best-of-N planning selects branches, while DiffusionNFT updates action parameters using the same scores. Separate studies examine full-history context and tactile refinement. The main physical-task suite reaches 84% average success after robot-domain mid-training; policy improvement and the other extensions are evaluated separately. [E02, E06, E13, E15–E17]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

DELE-w0.5: Inferring Action from Future Latent State for Robotic Manipulation

Lei, Fenghao; Huang, Zhixiong; Yang, Long; Chen, Jiabao; Huang, Peilin; Fu, Han; Li, Zhuo; Ren, Xiaoxue

DELE-w0.5 trains a manipulation policy with action-chunk generation and future visual-latent prediction, then removes future-observation tokens at deployment. A central tension is that the stated attention mask prevents actions from accessing predicted futures, despite the paper's future-to-action narrative. Its concrete deployment description therefore supports action generation with auxiliary future-state training rather than explicit inference from a generated future. The policy completes 50 of 80 physical trials across four fixed-scene tasks; the experiments do not isolate the contribution of future prediction. [E02, E05, E06, E09, E14]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

AcrossWAM1.0:A Modular Latent World-Action Stack for Compact Robot Policies

Zhang, Yafei; Wu, Nan

AcrossWAM1.0 studies whether LaWAM's latent-subgoal policy can support explicit backbone interfaces, a smaller multimodal backbone, and an auditable deployment export. A predicted transition latent is decoded into future visual features that condition continuous action generation. The strongest evidence is a paired LIBERO comparison: the compact policy achieves 97.45% versus 98.00% for the larger configuration. Cross-family evidence is limited to adapter execution, and inconsistent RoboTwin entries and parameter labels require clarification. [E02, E03, E09, E12, E13, E15]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Flex-ππ: A Multi-Stream World-Action Model with Compute Flexibility

Yan, Ge; Liu, Jinghao; Fan, Yuzhi; Cai, Lei; Liao, Minwen; Zhang, Jesse; Fox, Dieter

FLEX-π trains a 6B manipulation policy to predict future appearance, geometry, semantic features, and actions. Its distinctive mechanism separates the supervision available during training from the visual futures computed during deployment: every target stream trains on every sample, while attention masks teach the action expert to operate with different available streams. Action-only deployment can consequently retain benefits from geometric and semantic training without generating those futures. Joint generation remains useful on difficult manipulation tasks, but costs additional computation. The fastest reported configuration still consumes current multimodal observations; action-only does not mean RGB-only. [E02, E04, E05, E14, E22]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Hydra: A Navigation World Action Model with Discrete Latent Planning and Continuous Flow-Matching Execution

Nazeri, Mohammad; Card, Alexandyr; Huber, Samira; Pokhrel, Anuj; Wang, Yujun; Hammele, Ruben; Song, Daeun; Pirk, Sören; Xiao, Xuesu

Hydra learns a shared representation of navigation images, poses, and motor commands, then searches over quantized intents before producing continuous motion. Its central contribution is the combination of learned candidate generation and visual-latent safety scoring, which avoids rendering every imagined future. The physical experiments support efficient short-range planning on the tested setups, while long-distance failures expose limited goal-directed search convergence. The reported comparison used workstation inference, so onboard performance remains unestablished. [e02, e03, e07, e08, e12, e19]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models

Yang, Senqiao; Wang, Chengyao; Chen, Yuxin; Wang, Zixuan; Tang, Longxiang; Gui, Haokun; Ye, Jinhui; Lu, Changsheng; Wu, Xiaoyang; Zhu, Mingkang; Chen, Pengguang; Liu, Shu; Tian, Zhuotao; Zhao, Hengshuang; Yu, Bei; Jia, Jiaya

VLAct changes how a pretrained vision-language backbone learns from robot trajectories: protect early representations, train several continuous action decoders on shared features, and align compatible action coordinates across robots. The resulting backbone is transferred to a freshly initialized downstream action head. Controlled comparisons support better policy adaptation; an ablation excluding the downstream PI head provides narrower evidence for decoder-independent transfer. These are action-policy experiments, with no future-video prediction component. The supplied identifier, title, and authors match the catalog; the observed version is arXiv v1, dated August 27, 2026. [e01, e03, e04, e07, e20]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Making Latent Evolution Explicit: Operator-Structured Transitions for World Action Models

Lu, Xiaoxiao; Dong, Yunlong; Shi, Jiahao; Yuan, Ye

LEON replaces a latent future predictor's Transformer transition with context-modulated operators and additive forcing. It improves VLA-JEPA's LIBERO success and retains near-baseline LaWAM performance on four RoboTwin tasks. These are two different uses of prediction: training-time representation shaping and inference-time action conditioning, respectively (e02, e05, e06, e09, e10).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Riemann-1.0: An Embodied World Action Model for Physical AI

Sun, Haofeng; Pei, Jiangbo; Kang, Fei; Liu, Zexiang; Li, Yaokun; Jiang, Boyi; Xue, Hua; Zhou, Cindy; Li, Wei; Wei, Yichen; An, Mengyin; Zhao, Fanliang; Jiang, Biao; Wang, Zile; Liu, Yang; Li, Yangguang

Riemann-1.0 learns action prediction and visual dynamics in one transformer, ordering actions before their consequences. Robot deployment uses actual camera feedback; a second operating mode generates action-conditioned video. Pretraining moves from human-video pseudo actions through mixed trajectories to robot-only supervision. Strong reported task results accompany qualitative simulator examples, but no controlled ablation separates the effects of architecture, curriculum, and data scale. [E03, E04, E07, E08, E09, E11, E13, E16, E17]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization

Zhou, Jiaming; Zhang, Qihang; Xu, Gangwei; Fan, Cunxin; Zhao, Yujie; Wang, Ruilin; Luo, Yiming; Yang, Shuai; Zhu, Xing; Shen, Yujun; Liang, Junwei; Xu, Yinghao

Zero-WAM uses a human demonstration video to specify an unseen manipulation task, predicts the corresponding robot future, and decodes executable actions through inverse dynamics. Synthetic human–robot pairs and a training-only future-prediction objective support this interface. It reports 46.95% success on seven held-out RoboTwin tasks; real-robot evaluations test unseen configurations after embodiment adaptation. The evidence supports tabletop transfer under these protocols, with limited long-horizon reliability. [e02, e05, e06, e07, e08, e12, e13]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Tactile-WAM: Touch-Aware World Action Model with Tactile Asymmetric Attention

Wu, Siyu; You, Linjing; Zhu, Junjie; Liu, Yaozu; Kaixiang, Huang; Yonghang, Chen; Li, Jituo; Zhang, Changhao; Liu, Jian; Chu, Hengshuo; Li, Qi; Zhao, Hengshuang

Tactile-WAM jointly generates visual futures, tactile futures and action chunks in one Wan-based Transformer. Its directional attention mask protects the direct video pathway while an observed deformation cue strengthens action access to touch. Control gains are strongest relative to RGB-only DreamZero; visual-prior preservation and physical accuracy are distinct findings.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

4DGS-WAM: Bridging Past and Future with an Object-Centric World Action Model based on 4D Gaussian Splatting

Ma, Yueen; Xu, Zenglin; King, Irwin

4DGS-WAM reconstructs a persistent Gaussian scene, predicts actor motions, and transports existing object splats while retaining the static background. KITTI-MOT experiments support short-horizon image prediction and observed-view reconstruction; camera prediction and a frozen fusion underlay materially affect those results.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

SimWAM: A Simple World Action Model for End-to-End Autonomous Driving

Zhao, Zongchuang; Zhou, Xin; Xu, Tianyang; Sun, Zhengyang; Zhou, Kaixuan; Wu, Yu; Li, Honglin; Liang, Dingkang; Bai, Xiang

SimWAM trains separate video and trajectory diffusion transformers together, then plans from current observations without synthesizing future frames. Its strongest NAVSIM result includes action-expert reinforcement learning; additional experiments test imitation-only planning and open-loop transfer. The central evidence concerns efficient trajectory prediction, with limited support for broader deployment claims.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

From World Models to World Action Models: A Concise Tutorial for Robotics

Zhang, Xiaoxiong; Zeng, Xiong; Zhang, Wei

This tutorial separates predicting a robot's world from choosing actions within it. It organizes world models by prediction space, then compares four ways of connecting visual futures to policies. Its contribution is a conceptual design map, with qualitative tradeoffs and a notable training-versus-inference definition tension; it introduces no evaluated robot system.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

SANTS: A State-Adaptive Scheduler for World Action Models

Sun, Yirui; Zhuge, Guangyu; Liu, Keliang; Gu, Jie; Dai, Shiqin; Bing, Xinyu; Gan, Zhongxue; Tian, Chunxu

SANTS learns when to stop video denoising and how far to advance its noise level before a frozen action branch predicts a robot action chunk. Its reward measures demonstration-action agreement against shallow and full-denoising references, with a computation penalty. The v3 experiments report 94.4% RoboTwin 2.0 success at 523.7 ms and 73.1% mean real-robot success at 581.3 ms. These support a success–latency tradeoff within the evaluated tasks and backbone.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

GameWAM: A World Action Model for Video Games

Guo, Yuncheng; Zhang, Zhanqiu; Guo, Yiwen; Li, Weijia

GameWAM learns native game control through parallel video and action flow models, then deploys action-only denoising conditioned on learned visual context. It combines per-step gameplay/GUI routing, overlapping forecasts and bounded visual history. Results support the complete controller on Minecraft and ViZDoom; source interventions reveal that sampled noise can impose coherent camera bias (architecture, mask, mcu, vizdoom, lasi).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

GaussianWAM: Distilling Geometry and Semantics from 3D Gaussian Fields into World-Action Models

Zhang, Zijian; Jiang, Yuqing; Zhou, Weitao; Li, Minglei; Zhang, Jinhao; Mu, Yao; Li, Xiaofan; Zhao, Hao; Yu, Haibao

GaussianWAM strengthens existing world-action models by fitting an offline Gaussian field that binds estimated geometry to CLIP features, then teaching current-observation tokens to predict its rendered outputs. The field disappears at deployment. FastWAM's LIBERO-Plus success rises from 52.05% to 71.29%, although direct CLIP and VGGT distillation already reaches 69.37%; the distinctive contribution is therefore the incremental benefit of spatially organizing teacher signals. [e03, e05, e07, e12, e16]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

GeoWAM: Visual Geometry World Action Models for Autonomous Driving

Lu, Yiren; Ye, Xin; Liu, Jiaming; Jacobson, Philip; Yao, Jin; Chen, Yi-chung; Merino, Liam; Kurra, Dhruva Dixith; Cai, Min; Lampo, Tom; Yin, Yu; Guo, Danhua; Yaman, Burhan

GeoWAM forecasts dense future geometry from historical multiview images, then conditions a separate ego-motion decoder on predicted geometry tokens to regress one driving trajectory. Geometry pretraining and joint planning finetuning yield leading scores in the supplied NAVSIM tables, but the paper does not isolate which architectural component causes the gains.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

GlanceWAM: Sparse Test-Time Imagination for World-Action Models

Wang, Linhan; An, Zijian; Zhang, Mingyuan; Dai, Chen; Xu, Yi; Cui, Can; Yang, Zichong; Chen, Yinlin; Zhou, Lifeng; Lu, Chang-Tien

GlanceWAM uses one video transformer to imagine a future latent scene and condition faster action decoding on it. It reports 72.2% RoboCasa success and 99.0% LIBERO success; 48 ms measures an optimized LIBERO hold phase. Controlled ablations support the lookahead channel, while inconsistent refresh descriptions limit the deployment claim (e02, e09–e14).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

ReWorld: Representation Learning for World Action Models

Xia, Tianze; Zhou, Lijun; Xiong, Kaixin; Yao, Jingfeng; Zhu, Zhenxin; Sun, Haiyang; Wang, Bing; Chen, Guang; Liu, Wenyu; Ye, Hangjun; Wang, Xinggang

ReWorld trains the internal connection between a video generator and a trajectory planner. It supervises intermediate video states, aligns action states with attended video information, then fine-tunes both networks using nearby low-scoring trajectories. Reported gains cover generated video, simulated driving and frozen action-recognition features. The strongest video result combines representation training with changed sampling; the driving result concerns non-reactive NAVSIM simulation.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

LD4WAM: Learning Latent Dynamics from Human Videos for World Action Models

Shen, Zhenhao; Liang, Jiaqi; Lu, Jasper; Jiang, Feng; Wang, Yuran; Wei, Chuanbo; Liu, Jiayi; Yang, Jianchun; Yu, Qize; You, Jiadi; Hao, Ce; He, Guanqi; Xie, Chen; Wu, Ruihai

LD4WAM connects generated robot futures to action through a learned representation of motion. A separately trained Latent Dynamics Model supplies semantic transition targets grounded by end-effector supervision; a three-expert transformer then predicts video, distills those targets through learnable queries, and generates robot actions. The strongest mechanism evidence combines improved motion readout with executed-control ablations. Object transfer improves, but background robustness remains below the evaluated VLA baseline. [e-ldm, e-bridge, e-probe-results, e-ablation, e-generalization]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

WAM-OPD: On-Policy Distillation for World Action Models

Yang, Liuhaichen; Jiang, Zhuang; Sheng, Chenchao; Tang, Zezhi

WAM-OPD adapts an accelerated video-first robot policy using frozen-teacher labels on histories collected by the student. Actions are trained under the student's own predicted video, and shared adapters receive video, action and flow-matching losses. A small two-task simulation pilot improves success, while component causality and continuing on-policy adaptation remain untested.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

ForeTime-VLA: Causal Future-Token Distillation from a World Action Model for Conveyor-Belt Manipulation

Ma, Siyuan; Zhang, Yutian; Zhang, Boshi; Wu, Qinglian; Zhai, Jiaqi; Wei, Dong; Huang, Xiaojin

ForeTime-VLA teaches a π0.5 policy to predict a compact future code from causal history, then feeds that code into its vision-language and action paths. Privileged video supervises training; deployment produces action chunks without the teacher or future frames. Modest reconstruction improvements accompany larger real-robot grasp gains. The experiments support useful temporal conditioning, but do not isolate future distillation from added history, capacity and event supervision (e-overview, e-conditioning, e-offline, e-robot, e-speed, e-ablation).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

TacWAM: Anchor-Guided World Action Model with Mechanics-Aware Tactile Prediction

Jin, Lei; Ma, Yiding; Zhang, Xin; Gao, Chen; Wu, Wei; Li, Yong

TacWAM trains one visual–tactile–action generator to predict contact-relevant tactile latents alongside visual futures and robot action chunks. A mechanics-supervised encoder, recent tactile history, and restricted attention define how this supervision reaches the learned system. Actions read current anchors, not future sensory tokens. Four real-robot tasks yield 75.0% mean success versus VT-WAM’s 37.5%; staged ablations support the complete design but leave individual mechanisms only partly isolated.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

On the Capability Separation Between World-Model Policy Learning and Imitated World-Action Models

Yu, Yang

This theoretical paper asks whether predicting a future before decoding an action enlarges control capability. It separates representable closed-loop behavior, the target selected by ideal imitation training, and information about alternative actions. A world-action controller can be marginalized into a direct stochastic policy; under explicit population assumptions, both imitation objectives recover the same behavior-action conditional. Action-conditioned control additionally needs identified consequences and a utility-based decision rule. The strongest separation is a proved information gap in a two-environment construction, not a robot benchmark or an equality claim for finite neural architectures.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

WA-JEPA: Rethinking the Video JEPA Paradigm for World-Action Modeling in Autonomous Driving

Wang, Xinlin; Xiang, Yujiao; Zhou, Yuheng; Wang, Jingqi; Huang, Minqing; Huang, Jiajie; Wei, Dongxu; Zhou, Tingguang; Wang, Xiyang; Chen, Gong; Xu, Zhi; Tan, Feiyang; Zhou, Hangning; Yang, Mu

WA-JEPA turns video JEPA features into a driving planner through future-masked adaptation and joint flow generation of scene latents and ego trajectories. It reports strong NAVSIM and simulated closed-loop results, with evaluator-version and sampling-uncertainty qualifications detailed below.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

DECOWAM: Decoupled Whole-Body World-Action Model for Legged Mobile Manipulation

Ma, Siyuan; Zhang, Boshi; Zhang, Yutian; Wu, Qinglian; Zhai, Jiaqi; Wei, Dong; Yu, Qiaojun

DECOWAM adapts FastWAM to a moving quadruped–arm platform by separating arm control, base control, and camera ego-motion. It jointly predicts future video and actions using small trainable additions to an already adapted backbone. Replay errors improve, while hardware evidence favors coordination and displacement robustness more clearly than final task completion.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Towards Surgical World-Action Modeling: A Preliminary Joint Visual-Trajectory Forecasting for Surgical Motion Planning

Huang, Weiliang; Liu, Huanrong; Zhang, Bob; Dou, Qi; Chen, Zhen; Gu, Yun; Rosman, Guy; Li, Qingbiao

This preliminary surgical predictor uses historical video and tool locations to forecast both visual latents and image-plane trajectories. Repeated three-step prediction improves every reported metric over direct fifteen-step prediction, but error grows with horizon. The evidence establishes joint forecasting on recorded surgical sequences; it does not establish action-conditioned planning or successful robot execution.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

What Matters for Latent Actions in Robot Learning

Xizhou Bu; Qingda Hu; Lei Zhou; Lingfeng Zhang; Yingbo Tang; Zihao Liu; Xinyi Tao; Zhiqiang Ma; Qingqiu Huang; Chufeng Tang; Hongbo Wang; Jing Zhang; Jiayi Ma; Hangjun Ye; Wei Li; Xiaoshuai Hao

This empirical study asks which video-derived latent actions improve robot policies. It compares modeling paradigms, regularization, and action integration under a shared training framework. Raw-frame LAPO remains competitive, but preferred dimensionality and action heads depend on the evaluation. Its strongest deployment evidence is improved Franka manipulation after latent-action tuning of a VLM backbone; the forward predictor supplies training supervision rather than an inference-time planner.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

RISE: Adaptive Imagination for World Action Models

Lu, Hongbo; Yao, Liang; He, Chenghao; Han, Hao; Liu, Fan; Liao, Wenlong; He, Tao; Peng, Pai

RISE decides how long an autonomous-driving world-action model should imagine before producing a trajectory. A learned evaluator predicts prefix risk and continuation gains; a cost-conditioned gate repeatedly chooses Roll or Stop. CounterDrive supplies generated incident alternatives for prediction and localized risk supervision. RISE reports 90.8 EPDMS on NAVSIM v2 and a better score/latency combination than latent-convergence stopping. The evidence supports adaptive computation in the tested driving systems, while implementation conflicts and unverified deployment performance limit reproduction claims. [e02, e03, e05, e12, e14, e20]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

HiTac-WAM: A Hierarchical Tactile World Action Model for Contact-Rich Robot Manipulation

Xue, Chao; Zhang, Chaofan; Ma, Wenxuan; Yao, Guocai; Cui, Shaowei; Wang, Shuo

HiTac-WAM attaches an explicit contact–deformation–slip forecast to each action candidate generated by a pretrained world action model. Touch affects candidate selection and execution monitoring through a directed tactile branch. The strongest controlled comparison improves robot success using tactile ranking at a fixed candidate budget; forecast accuracy and anomaly detection retain important evaluation gaps.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

BrainWAM: Action-Space Coordination of Semantic Priors and Predictive Dynamics for Autonomous Driving

Zhan, Bing; Shang, Shuyao; Lu, Shuo; Xu, Yuan; Wang, Zhao; Wang, Yida; Zhang, Xueyang; Zhan, Kun; Gu, Jiahao

BrainWAM turns a vision-language planner and a video/action world model into two specialized action streams, then coordinates them through gated cross-attention and learned fusion. It improves reported NAVSIM planning scores while retaining costly video computation at inference; the evidence concerns non-reactive simulation, not physical vehicle deployment.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

WNM-3D: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN

Huang, Yuehao; Wu, Yunzi; Zhang, Xiaotao; Li, Xinhai; Dong, Jiankun; Lv, Jiajun; Zhang, Chi; Bai, Chenjia; Liu, Yong; Li, Xuelong

WNM-3D conditions a shared video–action diffusion transformer on geometry-aware tokens extracted from RGB history. It jointly predicts future visual latents and continuous navigation actions, then executes only the first action block before observing again. On GN-Bench, the full curriculum improves both Seen and Unseen success over WNM-2D. The evidence supports the combined geometry-conditioned design and staged training, while leaving component attribution, physical deployment and uncertainty unresolved.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Hydra-0: Action Flow for Generalist World Modeling and Control

Li, Hongyu; Wen, Bowen; Zhu, Xinghao; Wang, Yixuan; Du, Yilun; Li, Yunzhu; Konidaris, George; Birchfield, Stan; Pouya, Soha; Li, Chenran; Chang, Yan

Hydra-0 conditions video prediction on visible point trajectories, transporting initial-image features along their paths. Forward mode projects robot motion into this interface; inverse mode takes desired object motion and reads executable actions from denoised transformer features. Prediction and recorded-trajectory replay have quantitative support; physical control is illustrated on one pipe-bending task. [E03, E05, E08, E11, E15, E16]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Teach and Grow: An Agent-Centered Architecture for General Robot Learning

Nie, Chang; Liu, Zhe; Wang, Hesheng

Teach-and-Grow Learning (TGL) acquires executable, scoped Skill Blocks while leaving pretrained model weights fixed. An agent composes those blocks, checks physical effects and preserves lessons in memory. Small LIBERO studies support local acquisition and persistence; the paper's broader benchmark superiority and lifetime scaling predictions remain insufficiently quantified or untested.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

StageWAM: Joint-Embedding Stage Prediction for World-Action Models in Robot Manipulation

Liu, Xiao; Yang, Yuguang; Wang, Xi; Jiang, Kai; Chi, Cheng; Xu, Yong; Ding, Wenchao; Chen, Yilun; Wang, Yan

StageWAM predicts a latent next-stage target with Stage-JEPA, then conditions a Motus world-action model to generate local video and actions. It improves simulated manipulation success and reports shorter successful episodes. Its strongest ablation evidence concerns the extra conditioning pathway; the added benefit of task-specific stage training is small.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud

Chenghua Wang; Daliang Xu; Dongqi Cai; Duojin Sun; Hao Zhang; Haoze Qian; Huaiyuan Zhang; Jinshuo Cui; Junbo Cui; Kezhao Zhao; Longxi Gao; Mengwei Xu; Rongjie Yi; Ruixin Liu; Shangguang Wang; Tam Sikyuen; Tianyue Zhang; Weikai Xie; Xuanzhe Liu; Yingying Qin; Yiwen Lu; Yuan Yao; Yuezhi Zu; Yunhan Guo; Yuxin Zheng; Ziqi Guo

PhyAI is a shared inference runtime for existing vision-language-action and world-action policies. Model adapters retain policy semantics while runners, graph replay, fused kernels and parallel services accelerate execution. The experiments establish model-runner speedups and model-dependent batching behavior; appendices add simulation success checks. Control-rate ceilings and RL speedup projections require separate interpretation from measured latency.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

RoboSynChallenge: Mastering Real-World Dexterity via Generalizing Synthesized Manipulation Skills

Zhao, Runyi; Wu, Ruixin; Li, Chengkun; Zhang, Hongrui; Li, Ang; Jin, Ruixing; Deng, Yueci; Guo, Yingying; Ding, Lihe; Dong, Shaocong; Xue, Tianfan; Gao, Yanjun; Luo, Yudong; Poupart, Pascal; Wu, Simo; Jia, Kui; Zheng, Wei-shi; Liu, Guiliang

RoboSynChallenge proposes a benchmark for learning manipulation policies from synthesized trajectories and evaluating their transfer on physical dual-arm robots. Its contribution is a data-generation and evaluation protocol: EmbodiChain expands simulated experience, a smaller teleoperated dataset supplies real correspondence, and task success is checked through robot execution. Initial simulation-only and real-only baselines show task-dependent transfer, but inconsistent table aggregates and incomplete training specifications limit conclusions about data quality and generalization.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Keep the Future, Drop the Rollout: RIFT for World Action Models

Zhang, Chushan; Tong, Jinguang; Li, Xuesong; Wang, Yikai; Li, Hongdong

RIFT retains explicit future conditioning for robot actions while replacing iterative video generation with one learned cache prefill. Its central distinction is between consuming a future representation and producing it. Paired interventions motivate a fixed future interface; separate training and simulation experiments test whether anticipation tokens can supply that interface efficiently.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

4D-WAM: Infusing Spatiotemporal Awareness into World Action Models through Trajectory Fields

Yang, Lishan; Song, Wenxuan; Wang, Xi; Sheng, Pingyue; Fang, Zheng; Zhou, Ziyang; He, Junjie; Yan, Haodong; Chen, Jiayi; Sun, Nan; Sun, Qiao; Wang, Pengwei; Liu, Lingqiao; Wang, Yan; Gao, Yuxiang; Dayoub, Feras; Li, Haoang

4D-WAM trains existing video–action models with trajectory-informed auxiliary losses. It aligns adjacent-frame feature changes and first-to-final-frame correspondence distributions with a frozen Trace Anything teacher, then removes the teacher and projectors for deployment. The strongest evidence is improved simulated robustness under distribution shifts; standard benchmark gains are small, and real-world completion remains low. Several reporting inconsistencies limit precise interpretation.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Foresight Without Seeing: Latent Futures for World Action Models

Huang, Jiakai; Wu, Zhongbo; Zhang, Zheng; Wang, Zihan; You, Shan; Huang, Tao

ForeWAM gives a direct robot policy hidden predictive context through one Video DiT prefill, then reuses its K/V cache while an Action DiT denoises executable actions. A frozen teacher shapes dynamics registers only during training. Reported success is 96.7% on standard LIBERO and 61.6% on the observed LIBERO-Plus subset; Flash trades some subset robustness for lower latency. These are benchmark control results, with explicit coverage limits (e3, e5, e6, e9–e11).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Surgical WAM: A World-Action Model for Data-Efficient Surgical Robot Learning

Bao, Wenrui; Jiang, Tianyun; Chen, Zhiben; Lim, Ser-Nam; Peng, Peter D.; Shang, Yuzhang

Surgical WAM adapts Cosmos Policy so one diffusion transformer jointly predicts future endoscopic observations and robot actions. Surgical video pretraining precedes fine-tuning on a fixed action-labeled dataset. Receding-horizon execution converts these predictions into simulated dVRK control. The reported four-task mean improves from 63.5% to 77.8%, but incomplete protocols and inconsistent learning-efficiency presentations limit reproducibility; real-video evidence remains qualitative. [e-architecture, e-pretrain, e-control, e-main, e-curve, e-real]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

World Action Models in Real Time: An Empirical Study of Smooth Execution via Asynchronous Deployment

Team, Motubrain

This study compares six deployment strategies for a high-latency action-chunk policy on a physical bimanual robot. Learned prefix conditioning gives the strongest overall task-performance balance, while denoising-time action blending minimizes jerk but can impair insertion. Correct observation-to-command alignment remains fundamental. Offline chunk agreement and online milestone scores provide complementary evidence; neither establishes a new world-model architecture. [E2, E7, E8, E10, E11, E12]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

FACT: Failure-Aware Causal Training for World-Action Models

Peng, Quanquan; Liang, Yutong; Yan, Rui; Hansen, Nicklas; Wang, Xiaolong

FACT learns actions and their consequences in one video transformer. Successful demonstrations supervise all outputs; failed rollouts supervise future video and task progress while their action-imitation loss is disabled. Action-only deployment skips future prediction, with optional action-conditioned scoring offering additional gains at extra computation.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

4D-WAM: 4D Consistent World Modeling for Autonomous Driving

Fu, Jiacheng; Yuan, Yibo; Tian, Meng; Li, Yue; Zhu, Jiangtong; Han, Jianhua; Zhang, Yueyi; Fang, Jianwu; Xue, Jianru; Xu, Hang; Xiong, Zhiwei

4D-WAM jointly predicts driving video and ego trajectories, using a frozen geometry teacher during training to align generated futures with geometric responses to recorded futures. It emphasizes high-noise training timesteps and improves reported NAVSIM planning scores; the teacher adds no inference module. The evidence supports benchmark gains, while physical consistency and cross-domain robustness remain less directly quantified.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

SLIM-0.5B: Learning Action-Grounded Predictive Latents for Robot Manipulation

Wang, Jingkai; Tang, Zihan; Zhang, Gu; Cao, Mingyu; Chen, Jiapeng; Zhao, Jingjiao; Chen, Xiansheng; Wang, Pengwei; Liu, Lemao; Dou, Dejing

SLIM learns manipulation representations by connecting actions and observation changes in both directions. A shared observation–action transformer learns predictive latents before flow-matching policy training; deployment retains learned future slots without requiring future images. The 0.47B trainable policy combines strong benchmark performance with low measured inference cost, subject to differing baseline pretraining and native sampling configurations.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling

Tang, Qu; Zhuang, Benhui; Yuan, Bo; Yu, Xue; Guo, Longteng; Feng, Junlan

World Tokens makes a video-supervised representation the mandatory interface between a vision-language model and an action expert. Its current-context tokens support future-video denoising during training; deployment retains only the VLM, adapter, and action expert. Controlled LIBERO and physical R1 Pro results support improved control, with modest adapter overhead but substantial training cost.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

HarnessWAM: Bridging Prediction and Deliberation in World Action Models

Gu, Zhaopeng; Zhu, Bingke; Lin, Tianxi; Zhu, Guibo; Chen, Yingying; Wang, Kai; Yuan, Tingyu; Zhao, Chaoyang; Li, Zhaowen; Su, Peng; Wang, Jinqiao

HarnessWAM adds persistent scene evidence, constrained skill planning, event-triggered verification, and local recovery around LingBot-VA. Its matched-checkpoint comparisons support improved composition of manipulation skills; the strongest measured ablation concerns the interface that compiles semantic plans into executable operations.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Rethink Before You Execute: Adaptive Execution for World Action Models

Ye, Feng; Zhao, Yiming; Yu, Yong; Zhou, Hongxu; Pan, Yong; Xue, Yuan; Jia, Peng; Jia, Chuanmin

TempoWAM changes when a robot requests another action chunk. A recurrent monitor predicts progress after the next candidate actions; a calibrated protocol either executes them or requests replanning from a frozen WAM. Experiments show task-dependent computation savings or success gains, with normalized-time supervision limiting applicability to non-repetitive tasks.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling

Lin, Yihan; He, Jiawei; Bao, Shifeng; Zhao, Chen; Li, Yang; Wang, Xiaobo; Wang, Yan; Chi, Cheng; Zhang, Jing

JEPA-WAM couples dense current–future embedding prediction with robot action learning, then removes the transition head at deployment. Frozen V-JEPA features feed a Qwen predictor whose dedicated action representations condition a separate flow-matching expert. The experiments support improved scene-shift robustness, while exposing limits on language-dependent futures and the distinction between real-world completion scores and binary success.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models

He, Junjie; Li, Junfeng; Zhong, Zhide; Yan, Haodong; Li, Ruixin; Zheng, Yangyang; Zhu, Jiaguan; Zhang, Tianran; Du, Yuqiao; Chen, Wen; Zhou, Shunbo; Li, Haoang

SG-WAM adds a VLM planner that forecasts dense semantic and geometric features of future observations. These features condition a video expert, while an action expert receives their influence through joint attention. The reported advantage is small on standard LIBERO but larger under LIBERO-Plus perturbations. Physical trials support the approach separately from illustrative generated-video examples; the bar chart confirms the reported ordering but supplies no printed exact success rates. [e03, e06, e10, e11, e13, e14]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Vid2WAM: Distilling Video Diffusion Priors into World Action Models

Qiu, Chenhao; Wang, Ruixiang; Zhao, Runyi; Lin, Sixu; Gu, Songen; Nan, Shufeng; Liu, Guiliang; Jia, Kui; Fu, Yanwei; Wu, Simo

Vid2WAM trains a robot policy from both expert demonstrations and offline video-teacher rollouts, transferring generated futures directly into a student video objective and through inverse dynamics into action labels. Source-specific action adapters accommodate the two supervision sources. Deployment uses current-observation features and direct action prediction, without online video generation. RoboTwin transfer improves, but held-out LIBERO success remains much lower than its full-benchmark averages (e03–e12, e18).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

ω-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation

Li, Zhe; Zhang, Zhenzhe; Wei, Yangyang; Zhang, Wenjie; Yuan, Xichen; Zhi, Peiyuan; Li, Gen; Guo, Xinying; Gao, Fengjie; Yang, Jianfei; Zhang, Shanghang

ω-0 learns language-conditioned humanoid actions with auxiliary prediction of future visual latents. A pretrained action VLM and a joint query predictor condition a diffusion action head; SONIC executes its whole-body latents. Reported real-robot results favor this design, but camera differences, bundled ablations and inconsistent progress definitions limit causal interpretation.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots

Xu, Ling; Li, Borui; Wu, Hao; Han, Chuyu; Li, Xiangyu; Hua, Mohan; Jiang, Shiqi; Cao, Ting; Li, Chuanyou; Zhong, Sheng; Wang, Shuai

Embodied.cpp proposes a common C++ execution and deployment framework for VLA and world-action models, with pluggable heads and stateful scheduling. [e02, e07, e09] Its strongest reported VLA configuration, 4-bit HY-VLA, has latency and VRAM ratios of 0.37 and 0.23 versus Python, while retaining 0.99 of baseline success. [e13] The results expose useful deployment trade-offs, but the paper omits the task protocols, hardware and mechanism ablations needed to establish reproducible control performance. [e16]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Data Pyramid for Embodied Manipulation: A Survey

Yifan Ye; Yankai Fu; Yaoxu Lv; Bohan Hou; Jun Cen; Lingdong Kong; Duo Zheng; Tianxing Chen; Jiaming Liu; Ziang Cao; Yunfan Lou; Wei Chow; Xian Sun; Yingshuo Wang; Kuangzhi Ge; Xiaowei Chi; Xidong Zhang; Zhibo Pang; Yiwu Zhong; Sirui Han; Zhihe Lu; Weihao Yuan; Qifeng Chen; Michael Yu Wang; Yao Mu; Ziwei Liu; Jianfei Yang; Ping Luo; Shanghang Zhang

This survey organizes manipulation data by what becomes easier to collect and what remains grounded in a robot's actions. Its five-layer pyramid connects real-robot, UMI, human-video, simulation, and general data to embodied reasoning, action generation, and world prediction. The useful output is a framework for choosing and aligning supervision; it does not establish an optimal mixture or introduce a new policy. The authors explicitly warn that more sources and more trajectories need not produce better control.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Selective Cross-View Consistency for World Action Models: Held-Out Viewpoint Robustness Without Test-Time Camera Information

Huang, Bingqi; Wei, Bingchuan; Cai, Yingkai; Wang, Zhaokui

SCVC trains a joint video-action denoiser to agree across camera views on actions, future proprioception, and value, while preserving camera-specific future images. [E02, E03] Matched simulated experiments support extrapolation gains beyond trained camera ranges; interpolation does not improve and regresses significantly in the second seed. [E12, E13] Wrong-coordinate and wrist-masking diagnostics explain the objective and its evaluation contract. [E06, E10]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Decoupling Intention from Trajectory: A Representational Deduction Framework for World Action Models

Ma, Xiangkai; Ma, Yue; Wang, Junjie; Xu, Sheng; Li, Mingyang; Zhang, Han; Zhuang, Yuzheng; Li, Wenzhong; Yuan, Zhihao

PILOT trains motion-semantic query tokens to predict future VJEPA2-AC representations, then uses those tokens to guide action flow matching from current observations. Future-image generation and the Causal Dynamics Engine supervise training but are bypassed during deployment. Reported manipulation gains are promising, although conflicting implementation details, evaluation protocols and table arithmetic limit reproducibility.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models

Yan, Haodong; Li, Junfeng; He, Junjie; Zhong, Zhide; Yu, MingMing; Song, Wenxuan; Zhu, Jiaguan; Zheng, Yangyang; Du, Yuqiao; You, Jiadi; Cai, Yingjie; Yan, Xu; Zhao, Guanyi; Liu, Bingbing; Li, Haoang

Robust-WAM regularizes an existing video-generation WAM's action stream with future semantic targets while retaining its VAE-space video path. Temporally indexed learnable queries receive frozen DINOv3 supervision during training and remain as internal tokens at deployment. Controlled backbone comparisons show higher average OOD manipulation success; the evidence concerns executed policies, not merely plausible generated videos.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Adaptive-WAM: Quality-Guided Early-Exit Planning from Intermediate Video-Diffusion Features

Ang, Sining; Yang, Yuguang; Wang, Yan

Adaptive-WAM turns a video-diffusion backbone into a driving planner that stops when an intermediate trajectory looks adequate. Video prediction shapes training; deployment uses one conditional feature pass and a separate quality verifier. Its adaptive result is 90.79 PDMS at 170 ms, while the higher 92.6 score belongs to a different fixed-exit, 64-proposal system.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

MobileWAM: Bridging World Action Models to Mobile Manipulation with Chain-of-Foresight

Fan, Zehua; He, Junjie; Song, Wenxuan; Wang, Xi; Lyu, Wenqi; Zhao, Linge; Li, Fuhao; You, Zihan; Yang, Yifei; Xu, Kaiming; Jiang, Qi; Jiang, Yue; Li, Haoang; Chi, Cheng; Gao, Feng; Li, Bailin; Wang, Yan

MobileWAM couples a pretrained video transformer to a whole-body action expert. A recurrent future-latent loss shapes current-observation features during training; deployment discards future prediction and denoises actions against cached features. Softly mixed action experts support base–arm coordination. The strongest evidence is improved closed-loop simulation success; physical results remain modest and the input-state specification is inconsistent.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

DynamicWAM: Dual-Path Motion Conditioning for World-Action Models in Dynamic Manipulation

Lou, Yunfan; Gao, Hewen; Zhu, Xiyu; Qiao, Zhuoran; Han, Xuan; Yang, Yifan; Ye, Yifan; Yao, Boxian; Pang, Zhibo

DynamicWAM gives a coupled video/action model two views of recent motion: normalized optical-flow images preserve spatial structure, while numerical tokens restore displacement scale and timing. A distilled video expert and cached inference support deployment with asynchronous action chunks. Reported success is 38.2% on synchronous DOMINO and 46.67% across twelve physical-robot tasks, subject to the stated protocol and robustness boundaries (e03, e04, e05, e09, e10, e12, e14).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

DreamWAM: Beyond RGB Future Prediction for World Action Models

Yuan, Shanglin; Zhao, Weiheng; Shi, Xin; Jiang, Haoyi; Guo, Xianda; Liu, Liu; Liu, Wenyu; Sui, Wei; Wang, Xinggang

DreamWAM trains coupled video and action experts to anticipate appearance, motion, geometry and semantics, while retaining RGB-only deployment. Its strongest evidence is improved executed-task success under unseen perturbations, including when future video rollout is disabled. Routing heterogeneous targets matters as much as adding them.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Faster-WAM: Efficient Inference-Time Future Conditioning for Robust World Action Models

Zhao, Weiheng; Jiang, Haoyi; Shi, Xin; Liu, Liu; Huang, Fan; Su, Zhizhong; Sui, Wei; Wang, Xinggang

Faster-WAM generates a robot action chunk using future-aware features cached from one video-expert pass. SparseMoT limits where action layers read those features; Interval KV-Fusion combines information across video depths. The strongest evidence is improved manipulation under distribution shifts with lower measured inference latency, rather than improved decoded-video quality. The methods and comparisons below trace these claims to the inspected v1 paper.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

LiLa-WAM: Lightweight Latent Reasoning World-Action Model for Robotic Manipulation

Yang, Fan; Su, Yuting; Wang, Xiaobo; You, Yuncheng; Fan, Fugui; Wu, Yuting; Wu, Minghui; Zhao, Chenxu; Ning, JiaHong; Jing, Peiguang

LiLa-WAM trains action generation and future-feature prediction inside one compact transformer stream built on frozen DINOv3 features. A demonstration-derived Visual Transition Token selects the task. Its strongest evidence combines competitive simulation success with controlled foresight-loss ablations, including physical execution; it does not establish general-purpose future simulation or unseen-task instruction following.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

UniNav: A Unified World-Action Diffusion Model for Visual Navigation

Zhou, Changqing; Luo, Yueru; Jiang, Zeyu; Chen, Changhao

UniNav adapts a pretrained video transformer to jointly predict local waypoints, future visual latents and camera-geometry tokens. Geometry denoising also makes extra video-only data useful in the reported RECON ablation. A separately trained Fast variant removes future-video tokens at inference. The evidence establishes offline trajectory-error and latency trade-offs, while closed-loop goal-reaching remains untested in the supplied results. [e3, e5, e7, e8, e11, e13, e14, e15]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

CoWAM: Coordination Contracts for Selective Policy Intervention with WAMs

Liu, Shuaijun; Wen, Qifu; Hao, Shuyang; Luo, Qi; Zhang, Chenglong; You, Feiyang; Wu, Chengyu; Su, Ningxin

CoWAM uses typed coordination obligations and calibrated future evidence to decide whether an existing bimanual policy action should change. Its frozen proposer supplies alternatives; the selector cannot invent missing actions. Outcome-blind comparisons report better coordination selection and simulated closed-loop completion, while conservative gates limit unsupported overrides.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Faster-WAM: Do World Action Models Need Deep Action Modules?

Ma, Liheng; Yang, Rui Heng; Zhang, Zhanguang; Clemente, Mateo; Hu, Ziwen; Cao, Tongtong; Zhang, Yingxue

Faster-WAM couples a pretrained 30-layer video backbone to a single-layer action head through learned cross-layer key/value fusion and positional alignment. Deployment uses current-frame features and ten action-denoising steps without generating future video. Its reported 66.5 ms latency accompanies a LIBERO-Plus improvement over Fast-WAM, but RoboTwin success falls and two result descriptions conflict with printed figures or tables. [e-fusion, e-rope, e-inference, e-latency, e-shifts, e-robotwin, e-discrepancies]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space

Zhao, Ruiteng; Zhang, Zhengshen; Su, Yue; Wang, Wenshuo; Li, Jiahui; Yang, Zhiyuan; Tay, Francis E. H.; Jr., Marcelo H. Ang; Zhu, Haiyue

SG-WAM trains a compact manipulation policy to predict future states of its own dynamics tokens, conditioned on demonstrated actions. A slowly updated copy of the policy supplies future targets, while a frozen geometry teacher shapes current image tokens. Deployment retains the online VLM, dynamics tokens and flow-matching action expert; future prediction is training supervision. The strongest transfer result is 73.0% on LIBERO-Plus, subject to heterogeneous published baseline protocols.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

EndoWAM: A Grounded World-Action Model for Generalizable Endoscopic Navigation

Lin, Jinsong; Pan, Zikang; Liu, Wanhao; Ng, Chi Kit; Shao, Liangjing; Yu, Zihang; Wang, Ziyu; Wang, Yin; Wang, Jiaxi; Teoh, Jeremy Yuen-Chun; Xiong, Zhiyong; Gao, Huxin; Ren, Hongliang

EndoWAM turns intermediate video-model features into discrete endoscope action chunks. During training, reconstructing future target crops makes those features task-aware; deployment drops that reconstruction branch and avoids decoding future frames. The paper reports 80.2% success on three-stage physical-phantom navigation, versus 27.1% for its strongest baseline, and 7.5 Hz control. Its key evidence is executed phantom navigation and component ablations, with clinical transfer and detailed reproducibility left unresolved (architecture, navigation, ablation, efficiency, missing-appendices).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

SelfWAM: A Self-Grounded Unified World Action Model for Fast Robot Control

Pan, Bikang; Liu, Fan; Lu, Haotao; Wang, Jingya; Shi, Ye

SelfWAM co-trains a direct robot policy with action-conditioned RGB and robot-mask prediction. Its two-expert architecture isolates action denoising from clean demonstration targets, while future-video queries can access those targets. At deployment it generates actions without future video. The evidence supports stronger action sensitivity and modest aggregate policy gains, with important limits on held-out video generalization and physical-trial scale.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

FlowPilot: Real-Time World-Action Modeling for Agile UAV Navigation

Wang, Runqing; Yu, Ding; Min, Pengyuan; Zhang, Xinhong; Xiao, Wei; Hu, Yu; Chen, Jie; Zhang, Fu; Wang, Gang

FlowPilot couples future-depth and trajectory denoising for onboard quadrotor navigation. Its action is five free Bernstein control points, converted into a state-consistent reference for a separate tracking controller. Simulated comparisons and physical flights support the approach, while the strongest mechanism test freezes future-depth latents during inference. Image decoding is optional; this does not establish that latent depth computation can be removed. The reported latency requires careful accounting.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Disentangling Visuo-Tactile Foresight: Oracle-Guided Interface Discovery for World Action Models

Yao, Zihang; Ding, Chaoyue; Yu, Yingying

Oracle Visuo-Tactile Foresight (OVTF) studies how an action policy consumes successful future RGB and tactile observations. Privileged reference trajectories replace a learned future predictor. Asymmetric Phase-Local Future Memory (AFM) compresses these streams into slots with selective visual-to-tactile attention. Across seven UniVTAC simulation tasks, AFM records 32.0% success versus 23.7% with isolated modality routing. This diagnoses an interface under oracle access; it does not establish performance with predicted futures.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Diagnosing and Mitigating Perception-Decision Misalignment in Omni-LLMs via Modality Subspace Activation

Jiang, Hongbo; Li, Jie; Shen, Yunhang; Xie, Tianyu; Dai, Pingyang

MSA intervenes in an Omni-LLM’s final hidden state using modality-specific SVD subspaces. Its companion benchmark and CMS metric test whether answers and option logits react when audio or vision disappears. All seven tested variants gain sensitivity, but only two gain accuracy; the central tradeoff is between responsiveness and correctness, not demonstrated robot-control reliability.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

FBFM: A Training-Free Asynchronous Feedback Mechanism for Flow-Matching in World-Action Models Execution

Li, Peize; Zhang, Ruimeng; Zhang, Ru; Huang, Cong; Chen, Kai; Zhang, Shanghang

FBFM inserts observation feedback into a frozen world-action model's active flow solver. It combines changing state measurements with fixed commitments from the preceding action chunk, improving selected RoboTwin task averages and producing a small, heterogeneous LIBERO gain. Its strongest contribution is the explicit correction path; deterministic timing and recorded-video diagnostics leave real-time physical control unestablished.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Paired Exact-Reset Evaluation of a Prediction-Derived Medium-to-Full World-Model Cascade

Pastor, Malo de

This paper evaluates when a Medium predictor's imagined consequences help select between separately frozen Medium and Full computations. Exact simulator resets supply paired physical decision losses for identical candidate actions. A prospective PushT confirmation finds lower latency-priced regret than a dimension-matched current-DINO/action router, although prediction routing is slower. The contribution is an evaluation protocol with explicit cost accounting and clustered inference; its evidence concerns controlled candidate-set decisions, without establishing closed-loop robotics value.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts

Wang, Mingxin; Hu, Bin; Qian, Bin; Jiang, Kaitao; Wu, Haoning; Yan, Feng; Jing, Bowen; Hao, Ruiyang; Wang, Enyi; Niu, Kangning; Yang, Yandan; Xu, Mu; Wang, Yan; Liu, Houde; Li, Tianlun

ST-WAM trains separate visual-future, semantic-future and action experts together, then deploys an action-only policy. DINOv3 supplies both semantic future targets and recent-history features retrieved under current image/language context. The central result is improved manipulation under visual shifts, supported by component ablations; it does not establish that correcting generated videos is necessary for control.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

QuantWAMs: Calibrating at the Right Granularity for World Action Models

Zhou, Jiacheng; Lv, Jinfan; Li, Ruixuan; Zhang, Longtai; Wang, Yan; Zhang, Wenqiang; Qi, Lizhe

QuantWAMs compresses existing world action models by matching activation pooling, weight precision and denoising-step protection to calibration evidence. On two WAMs, its mixed W4A4-dominant configurations report simulation means 0.2–0.7 percentage points below FP16, with lower block memory and latency. Additional gradients and FP16 rollouts are required; neither end-to-end acceleration nor statistical equivalence is established.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

EgoGenesis: Egocentric World-Action Modeling with Online Anchored Projective Memory and Action-3D RoPE

Yan, Zexuan; Wu, Yuzhou; Ma, Yue; He, Zonghang; Yin, Kaibo; Tu, Xiaobing; Wang, Yinggui; Ren, Jinkui; Zhang, Xiantao; Wang, Shijian; Liu, Jinghong; Zhang, Linfeng

EgoGenesis resimulates manipulation demonstrations with an autoregressive video generator conditioned on camera and action geometry. A permanent scene anchor and refreshed recent memory stabilize the scene; metric rotary attention aligns skeleton control. Synthetic videos improve a separately trained robot model, but video plausibility and executed success remain distinct evidence streams (architecture, oapm, a3d, robot-results, downstream).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

LeapBot-WA: World-Anchor Action Models via Predictive Latent Alignments

Liu, Pei; Zheng, Nan; Zhang, Lang; Peng, Daojie; Zhang, Yanan; Kong, Feilong; Feng, Mingyue; Liu, Jiachao; Wang, Yaonong; Chen, Qifeng; Ma, Jun

LeapBot-WA trains a manipulation policy with future semantic guidance, then executes actions using a cache of the current observation. Its V-JEPA features pass through an isotropy-regularized semantic autoencoder before asymmetric world/action training. Reported simulation performance is competitive, especially against the listed latent WAMs, but the supplied v2 contains unresolved objective, pretraining and numerical inconsistencies. The architectural idea is clearer than the exact reproducible recipe. [e-overview, e-cache, e-isae, e-libero, e-objective-conflict]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

DC-WAM: Dynamic-Centric Visual Supervision and Reasoning for World-Action Models

Ji, Haoyuan; Fan, Lingxiang; Su, Shang; Lu, Yinqiao; Shi, Mengkai; Gao, Jun; Feng, Shuo

DC-WAM trains a robot action policy with future-video supervision concentrated on temporal changes and tracked interaction regions. Its two-branch Wan2.2-based model adds DynaRoute attention biases, then deploys with one visual-cache prefill followed by action-only denoising. The clearest benefit is improved success under distribution shift; the supplied paper leaves important routing and reproduction details unresolved.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone

AI, Simple; :; Wei, Yuteng; Ma, Jinming; Wang, Jiawei; Zhou, Weitao; Zuo, Yushen; Rui, Ke; Li, Minglei; Zhang, Jinhao; Pan, Zhikang; Wang, Xiang; Jia, Haoran; Du, Huan; Zeng, Zicheng; Ma, Jun; Qin, Guiyu; Zhang, Di; Li, Xiaofei

HiFi-UMI combines wearable robot-free capture with offline reconstruction, simulation replay, and curation to produce action-grounded demonstrations. Across three policy backbones, task-specific UMI post-training reaches approximately the aggregate success of teleoperation under the tested collection regimes. This supports a practical data pipeline, with substantially more UMI demonstrations and matched grippers and wrist cameras.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

RoboHarness: Memory-Driven Orchestration of Heterogeneous Robot Policies for Long-Horizon Planning

Huang, Jinbang; Hu, Yuanzhao; Li, Zhiyuan; Qi, Ran; Xiao, Yixin; Zhang, Zhanguang; Coates, Mark; Cao, Tongtong; Zhang, Yingxue

RoboHarness routes subtasks among independently developed robot policies, then uses retrieved execution states to prepare compatible handoffs. It combines a coding-agent planner, capability assessments and online adaptation. The central empirical evidence is improved executed long-horizon manipulation and a substantial completion loss when bridging is removed; this is an orchestration method, not a new joint future/action predictor (e03, e06, e11, e13).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

N₀-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation

Team, NeoteAI; Team, Fudan TEAI

N₀-TWAM generates future camera and tactile observations, then conditions robot actions on that forecast and current touch. Its integrated transformer cascade improves average contact-rich task success, but transfer varies across tasks and sensors. [e-cascade, e-observed, e-real, e-generalization]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

GeoWorldAD: Geometry World Action Model for Autonomous Driving

Zhang, Songyan; Tian, Jinyuan; Li, Hanbing; Liu, Daqi; Chen, Hao; Huang, Wenhui; Li, Fang; Chen, Guang; Ye, Hangjun; Chen, Long; Yang, Kuiyuan; Lv, Chen

GeoWorldAD plans driving trajectories from video using ego-aligned present geometry and learned future geometry tokens. Its planner progressively refines proposals across geometry scales, then incorporates anticipated scene evolution. The most persuasive reported benefit is greater ego progress over the present-only GeoAD baseline. Evaluation concerns NAVSIM simulation; reconstruction quality and qualitative future depth provide supporting evidence, with important gaps in causal isolation and deployment detail.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

NVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulation

NVIDIA; Aarti Basant; Amlan Kar; Despoina Paschalidou; Fangyin Wei; Francesco Ferroni; Guillermo Garcia Cobo; Haithem Turki; Huan Ling; Jaewoo Seo; James Lucas; Jay Zhangjie Wu; Jialiang Wang; Jonathan Lorraine; Jun Gao; Kai He; Katarina Tothova; Kevin Xie; Michał Tyszkiewicz; Qi Wu; Riccardo de Lutio; Ruilong Li; Sanja Fidler; Seung Wook Kim; Tianchang Shen; Tianshi Cao; Tobias Pfaff; William Lew; Xindi Wu; Xuanchi Ren; Yifan Lu; Yuxuan Zhang; Zan Gojcic; Zian Wang

OmniDreams adapts Cosmos into a causal, action-conditioned camera simulator with few-step diffusion and reusable visual history. Its strongest systems result is four-camera rendering at 105 effective FPS per camera on sixteen GB300 GPUs. A separately post-trained trajectory policy also improves simulated collision rates. These findings establish rendering and policy capabilities under distinct protocols, with compute, decoder quality and chunk-level feedback tradeoffs.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

PerceptDrive: Perception Prior World-Action Modeling with Adaptive Expert Routing for End-to-End Autonomous Driving

Liu, Yushan; Lv, Tianxiong; Wang, Bohua; Fan, Hangqi; Zhao, Chenxu; Zheng, He; Zhong, Xuchang; Xie, Yifan; Zhao, Congyang; Liao, Zhihao; Luo, Leigang; Cai, Yang; Zhang, Xiao-Ping; Ding, Wenbo

PerceptDrive converts frozen geometric, semantic, and dynamic perception priors into driving trajectories. Branch-specific reconstruction preserves these priors through query compression; scene-dependent gates combine their conditions before a predicted latent future guides a flow actor. Its strongest evidence concerns NAVSIM planning scores and internal controls; interactive driving and transfer beyond the training evaluator remain open.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

WorldScape Policy 2.0: Empowering Steerable World Action Modeling with Reasoning-Augmented Memory

Su, Haisheng; Liu, Zongdai; Jin, Xin; Dou, Haoxuan; Hu, Chengming; Li, Baorun; Liu, Zhanwang; Xu, Ruiyan; Fang, Jianjie; Zhang, Xin; Yang, Zhenjie; Yang, Xue; Gao, Chen; Yan, Junchi; Li, Yong; Wu, Wei

WorldScape Policy 2.0 combines a shared video-action diffusion transformer with recent visual history and retrieved event memory. Event captions directly steer fine-grained actions or supervise an autonomous latent subgoal pathway. The paper reports strong RoboTwin and PiPER execution results, but its headline simulation score uses randomized training; clean-only generalization is substantially weaker. The evidence below separates these protocols and the two instruction modes.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Test-Time Scaling for World Action Models via Zero-Shot Geometric Evaluation

Zhao, Zesen; Cho, Minkyoung; shen, Hui; Zheng, Boyuan; Gao, Kunxiao; Cao, Yulong; Mao, Z. Morley

Gated GeoBoN uses a WAM’s predicted futures twice: inexpensive action–future agreement determines whether to draw more candidates, then geometric consistency ranks those candidates. It requires no new training. Fixed-budget selection improves all five reported benchmark–backbone averages at N=8; gating retains much of that benefit with fewer sampling decisions, while evaluator outliers limit larger budgets. [e-purpose, e-fixed-budget, e-gated, e-failure]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch

GigaWorld Team; Angen Ye; Angyuan Ma; Boyuan Wang; Chaojun Ni; Fangzheng Ye; Guan Huang; Guo Li; Guosheng Zhao; Haodong Yan; Hengtao Li; Jiwen Lu; Kai Wang; Mingming Yu; Qitang Hu; Qiuping Deng; Songling Liu; Xiaoyu Tian; Xiaofeng Wang; Xinyu Zhou; Xiuwei Xu; Xinze Chen; Yang Wang; Yejun Zeng; Yifan Chang; Yun Ye; Zhenyu Wu; Zhanqian Wu; Zheng Zhu

GigaWorld-Policy-0.5 trains action prediction together with future visual dynamics, then omits future-video tokens during robot deployment. Specialized visual and action Transformer experts, mixed pretraining and runtime optimization yield stronger reported manipulation scores and an 85 ms RTX 4090 C++ inference path. The evidence is real-robot execution, but small trial counts, partial-credit metrics and incomplete implementation details limit broader conclusions.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

BadWAM: When World-Action Models Dream Right but Act Wrong

Li, Qi; Yang, Xingyi; Wang, Xinchao

BadWAM tests whether bounded visual perturbations can separate a world-action model’s actions from its predicted future. It optimizes action deviation through queries, optionally penalizing future drift, then evaluates closed-loop control. LIBERO shows substantial reliability loss; closeness to clean imagination remains a limited proxy for stealth.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

AeroAct: Action-Centered World-Action Models for Language-Conditioned Quadrotor Flight

Zhang, Xinhong; Zhu, Qiyuan; Huang, Yubo; Chen, Haolin; Wang, Runqing; Mo, Yuhao; Chen, Zhongxin; Hu, Yu; Wang, Xinjiang; Sun, Jian; Wang, Gang

AeroAct adapts a video diffusion Transformer into a language-conditioned quadrotor policy. It learns actions alongside their future visual consequences, then deploys only the action stream. Its strongest controlled evidence concerns temporal visual history; the physical demonstration establishes execution in a short indoor flight setting. Evidence: E04, E13, E16.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Steering Robustness into World Action Models via Mechanistic Interpretability and Optimal Control

Hong, Jihoon; Skifstad, Julian; Dai, Qiyue; Chan, Alice; Chou, Glen

WA-LQR modifies pretrained world-action-model activations during inference to improve simulated manipulation under camera, gripper and image-noise perturbations. Contrastive examples identify nuisance-related directions; local activation dynamics support feedback corrections along those directions. Success depends on model and task geometry. Cosmos-Policy benefits most consistently on camera and gripper shifts, while open-loop ActAdd wins its Gaussian-noise average. LingBot-VA shows weak overall gains. These findings support selective activation steering, with calibration and transfer limits still unresolved.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

FlowWAM: Optical Flow as a Unified Action Representation for World Action Models

Chen, Yixiang; Li, Peiyan; Xu, Yuan; Ma, Qisen; Yang, Jiabing; Wang, Kai; Yang, Jianhua; An, Dong; Guan, He; Liu, Gaoteng; Si, Jianlou; Huang, Jun; Liu, Jing; Liu, Nianfeng; Huang, Yan; Wang, Liang

FlowWAM generates RGB futures and optical-flow plans with a shared video transformer, then decodes their hidden states into robot actions. Providing flow instead enables controlled video prediction. Its strongest evidence combines manipulation success, improved trajectory scores and representation ablations, with important preprocessing and evaluation qualifications.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence

Liang, Yuanzhi; Zhan, Xufeng; Huang, Haibin; Zhang, Chi; Li, Xuelong

This survey proposes explicit contracts linking physical prediction, model intent, execution and reusable experience. A World Action Model estimates intervention consequences; an embodied brain forms an intended change; a physical harness translates and verifies execution. The contribution is a research roadmap with testable responsibilities, rather than a trained embodied-brain system or demonstrated performance gain. Its promised reuse remains a hypothesis (e-contract, e-ownership, e-conclusion).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Towards Predictive, Aligned, and Scalable Robot Learning

Tang, Peijun; Xie, Shangjin; Huang, Baifu; Sun, Binyan; Yang, Haotian; Luo, Kuncheng; Jin, Weiqi; Fang, Shilin; Wang, Jianan

Lumo-2 trains a Qwen3.5-4B-based robot policy to predict compact visual dynamics before generating semantic action tokens. Progressive dynamics–action and vision–language alignment structures its action representation, while historical actions disambiguate execution phases. Real-robot results favor the system, but partial-progress metrics and bundled training changes limit causal conclusions.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Artificial Foveated Perception for Mitigating Shortcut Learning in Robotic Foundation Models

Sun, Xiatao; Zhuang, Yuan; Negrete, Mateo Sanchez Lopez; Coldea, Matei-Victor; Liang, Chen; Zhang, Haoyang; Liu, Che; Zeng, Ziyao; Li, Shawn; Wang, Qian; Miao, Fei; Rakita, Daniel

Artificial Foveated Perception (AFP) learns language-conditioned relevance masks and uses them to supervise a robot policy’s visual attention during fine-tuning. A protected auxiliary gradient improves reported distractor robustness across four simulation backbones and one physical robot policy. AFP is absent from the deployed control loop; annotation requirements and incomplete experimental specifications limit reproducibility.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

DIM-WAM: World-Action Modeling with Diverse Historical Event Memory

Wang, Kai; Gu, Zhaopeng; Chen, Yixiang; Xu, Yuan; Ma, Qisen; Yang, Jiabing; Li, Zhaowen; Huang, Yan; Wang, Liang; Su, Peng

DiM-WAM augments LingBot-VA with bounded, observation-grounded event memory that conditions future-video and action denoising. Bank-specific summaries preserve cross-stage evidence, while auxiliary trajectory-progress supervision shapes memory. The inspected v2 reports RMBench success of 69.8% versus 34.8% in a training-matched comparison, and 90.0% versus 52.5% full-task success on four Franka tasks. The combined design is supported; stable semantic bank specialization remains unestablished.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

FlowDAgger: Human-in-the-Loop Adaptation of Generative Robot Policies in Latent Space

Murray, Michael; Chen, Daphne; Bagaria, Simran; Fortier, Dean; Hellebrekers, Tess; Mullins, Galen; Gajarla, Harshavardhan; Mees, Oier; Cakmak, Maya; Kolobov, Andrey

FlowDAgger converts expert corrections into noise targets for a small policy that steers a frozen generative robot controller. It supports action-head models and joint world-action diffusion. Experiments report efficient adaptation and better retention than the tested alternatives, with meaningful residual forgetting and limited implementation disclosure.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents

Yixian Zhang; Huanming Zhang; Feng Gao; Xiao Li; Zhihao Liu; Chunyang Zhu; Jiaxing Qiu; Yuchen Yan; Jiyuan Liu; Wenhao Tang; Zhengru Fang; Yi Nie; Changxu Wei; Yu Wang; Wenbo Ding; Chao Yu

Harness VLA makes a frozen visuomotor policy a callable contact primitive inside a memory-guided planner. RGB-D grounding and analytic movement connect local VLA attempts; stored strategies guide recovery across changed scenes. Its strongest evidence is improved simulated task success with the same frozen backend. Interpretation depends on memory access, unequal external-baseline coverage, and a figure/table discrepancy. [execution-loop, bootstrap-memory, pro-results, overview-figure]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Understanding and Mitigating the Video-Action Generalization Gap via Temporal Ratio

Mishra, Utkarsh A.; Chen, Yongxin; Xu, Danfei; Liu, Yang; Chen, Xi; Mao, Jiayuan

Temporal Ratio measures how much a video-conditioned action head attends to imagined futures relative to the current observation. The paper uses this signal to schedule language and long-horizon guidance during action sampling. LIBERO and bimanual experiments support improved compositional control, with remaining dependence on video accuracy and unresolved implementation inconsistencies.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time

Feng, Yusen; Han, Bingchen; Lyu, Jiangran; Liu, Kai; Zheng, Yixin; Wan, Yuxuan; Liu, Weiheng; Han, Sun; Li, Ruiqin; Zhang, Yulong; Liu, Fangfu; Shi, Xuesong; Liu, Libin; Wang, Yizhou; Zhang, Zhizheng; Wang, He

WAM-TTT learns a memory interface that lets action-free human videos steer an LDA world-action model. Deployment adapts only video-side fast weights, then fixes that memory for robot execution. Reported household progress improves, but scene-exposure and loss-description inconsistencies limit the generalization and reproduction claims (e03–e06, e09, e15, e22).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

EgoWAM: World Action Models Beyond Pixels with In-the-Wild Egocentric Human Data

Li, Baoyu; Yin, Xinchen; Lin, Mengying; Zhang, Yixin; Xu, Danfei

EgoWAM studies which auxiliary future-prediction target helps a robot learn from egocentric human demonstrations. A shared transformer predicts actions alongside pixel latents, DINO features or stabilized 3D motion. DINO and 3D Flow produce stronger transfer than pixel prediction in the reported tasks. The world head shapes training representations and is discarded for deployment; the policy does not plan through generated futures.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

WAM4D: Fast 4D World Action Model via Spatial Register Tokens

Li, Ying; Wei, Xiaobao; Cao, Jiajun; Wang, Hao; Chi, Xiaowei; Bai, Chengyu; Sun, Qianpu; Li, Jiajun; Zhang, Xiaojie; Jia, Peidong; Tang, Jian; Han, Sirui; Zhang, Shanghang

WAM4D trains a causal video-action model to expose future geometry through auxiliary spatial registers, then deletes the depth branch for control. Its strongest mechanism evidence is the benefit of a trainable pretrained depth head; full-suite results show competitive success and reduced memory, with mixed accuracy and latency advantages.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Learning 4D Geometric Priors for Inference-Efficient World Action Models

Zhang, Jianjun; Zhu, Jian; Su, Taiyi; Ma, Chong; Huang, Zitai; Xu, Yi; Wang, Hanli

MECo-WAM adds a training-only geometry expert to a video-action model, then removes it for action inference. Frozen VGGT features supervise action-weighted spatial relations and temporal changes; temporary attention to current geometry decays away. Reported gains are modest in simulation and task-dependent on a real robot, with essentially unchanged action-chunk latency (e-architecture, e-attention, e-distillation, e-robotwin, e-real, e-latency).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

UNIVERSE: Unified Video Action Models for Autonomous Driving with Flexible Mask-Modulated Modality Generation

Liu, Mengmeng; Zhang, Diankun; Liu, Jiuming; Cui, Jianfeng; Xie, Hongwei; Chen, Guang; Ye, Hangjun; Nex, Francesco; Cheng, Hao; Yang, Michael Ying

UNIVERSE trains future driving video and ego trajectories in one diffusion transformer while blocking attention between their future tokens. Video supervision updates the planner’s parameters, but trajectory-only deployment omits video generation. Reported benchmark gains support this design without establishing physical driving safety (e03, e04, e05, e10, e11).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

DSWAM: A Dual-System World Action Foundation Model for Fine-Grained Robot Manipulation

Zhu, Jian; Zhang, Jianjun; Su, Taiyi; Liu, Tianbin; Wang, Zhangyuan; Xie, Kai; Huang, Zitai; Ma, Chong; He, Youzhang; Wang, Tianjian; Wang, Hanyang; Ding, Weihao; Xu, Yi

DSWAM trains a video-based robot executor with action and future-video supervision, then generates only action chunks during deployment. An optional language planner supplies subtasks. Its strongest physical comparison uses the executor alone: reported folding success is 96.3% versus DeMaVLA’s 92.5%, with timeout-inclusive completion time falling from 2 min 18 s to 1 min 44 s. The evidence supports the deployed system, while leaving the separate causal contribution of video supervision unresolved.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

ABot-M0.5: Unified Mobility-and-Manipulation World Action Model

Chen, Ronghan; Yang, Yandan; Tang, Zuojin; Huo, Dongjie; Lin, Tong; Wu, Haoning; Liu, Haoyun; Chen, Yuzhi; Zheng, Lulu; Yuan, Botai; Li, Tianlun; Wang, Mingxin; Qi, Dekang; Hu, Bin; Mei, Wei; Xuan, Yuze; Yang, Haolong; Zhu, Yanqing; Xu, Mu; Ma, Zhiheng; Chang, Xinyuan

ABot-M0.5 turns predicted video into frame-level motion representations and then robot controls, with separate mobility and manipulation branches. Dream Forcing trains the action predictor on its own model's imperfect futures. RoboCasa365 Target 100% success reaches 54.2%, but unseen-composition performance depends strongly on the protocol, and several training details remain inconsistent or unspecified. [e-cascade, e-mot, e-dream, e-target, e-pretraining-results, e-ambiguities]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies

Tianxing Chen; Yue Chen; Zixuan Li; Junyuan Tang; Kailun Su; Haoran Lu; Weijie Wan; Baijun Chen; Songling Liu; Haowen Yan; Honghao Su; Zhiyang Dou; Kaixuan Wang; Dandan Zhang; Yunze Liu; Yan Qin; Qiwei Liang; Qiwei Wu; Zijian Lin; Wenwei Lin; Yuran Wang; Minghua He; Tianshu Wu; Ruihai Wu; Jingquan Zhou; Kai-Chong Lei; Haibao Yu; Yuanfeng Ji; Weiyang Jin; Guanyu Lin; Xiaofan Li; Qi Xiong; Renjing Xu; Zhongyu Li; Wenhao Chai; Enze Xie; Ziwei Wang; Yao Mu; Hao Dong; Wojciech Matusik; Mingyu Ding; Wenbo Ding; Ping Luo; Masayoshi Tomizuka

RoboDojo couples 42 simulation tasks with 18 physical tasks through XPolicyLab and standardized RealEval hardware. Its contribution is an evaluation system: capability-specific scores expose brittle robot execution, while parallel simulation reduces measurement cost. The suites are complementary, not paired transfer tests.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

HALO-WA: Hybrid-Attention Latent-Guided Online Reinforcement Learning for World-Action Models

Ye, Angen; Ke, Weijie; Wang, Xiaofeng; Chen, Xinze; Ni, Chaojun; Zhao, Guosheng; Wang, Boyuan; Zhu, Zheng; Xie, Junjie; Zhang, Dapeng

HALO-WA uses a frozen world-action policy twice: its action chunk supplies a behavioral prior, and its distributed visual latents supply information for correction. A separate actor-critic learns refined chunks from real interaction. Four physical tasks show strong task-specific gains, supported by narrower ablations and simulation experiments; general-purpose transfer is untested.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Kairos: A Regret-Aware Native World-Action Model Stack for Physical AI

Team, Kairos; Wang, Fei; You, Shan; Zhang, Qiming; Huang, Tao; Fu, Zuoyi; Zheng, Zhisheng; Xi, Yunlong; Lv, Feng; Wu, Xiaoming; Liu, Zeyu; Wan, Cong; Li, Pu; Yang, Ruiqing; Li, Xiaoou; Wang, Wei; Zhu, Kangkang; Zhang, Yuwei; Fu, Shi; Zhang, Zheng; Wu, Xiaoning; Fan, Xuzeng; Tao, Dacheng; Wang, Xiaogang

Kairos couples a video diffusion transformer with an action transformer, sharing visual history and language grounding. Joint video-action training improves manipulation results, while hybrid attention supports longer video histories and action-only inference avoids generating future video. The evidence establishes benchmark capabilities and generation efficiency; it does not measure representation-induced regret or validate a closed-loop robot learning system.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

VT-WAM: Visual-Tactile World Action Model for Contact-Rich Manipulation

Tian, Shuai; Zheng, Yupeng; Zheng, Yuhang; Gu, Songen; Zang, Yujie; Qin, Yuxing; Li, Weize; Li, Haoran; Ding, Wenchao; Zhao, Dongbin

VT-WAM learns visual futures, tactile deformation and robot actions in an integrated multi-expert flow model. During control it keeps a current visual anchor and predicts tactile evolution alongside actions. A contact-gated training loss encourages action queries to use touch. Its six-task average is 71.67% versus Fast-WAM’s 45.00%, but surface-task scores include partial credit. Evidence supports task-specialized contact manipulation, with missing latency and uncertainty measurements limiting broader claims.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Bridge-WA: Predicting Where and How the World Changes for Robotic Action

Bai, Yongjie; Wang, Hanting; Dai, Mingtong; Zhong, Qijun; Liu, Yang; Lin, Liang

Bridge-WA compresses a manipulation-trained future teacher into outcome tokens, spatial change maps and motion-flow maps. A policy predicts these priors from current context and reads them while generating action chunks. Reported gains concern executed manipulation, with improved averages but uneven transfer across perturbations and tasks. Efficiency is architectural; deployment latency is not measured (e03, e10, e12, e13, e20).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Efficient Sim-to-Real Transfer of World-Action Models from Synthetic Priors

Wang, Zixing; Sivakumar, Kausik; Shang, Jinghuan; Hu, Yafei; Xie, Zhaoming; Gong, Ran; Zhang, Xiaohan; Schmeckpeper, Karl

This early-results paper adapts Cosmos Policy using automatically generated simulation demonstrations and deploys it directly on a Franka Research 3. Approximately 800 synthetic demonstrations per task yield 35% average success over four real manipulation tasks, without real policy demonstrations or real-data fine-tuning. Predicted/live image comparisons and an unseen-bottle rollout provide qualitative diagnostics. The experiments establish transfer feasibility in this setup; they do not isolate the contribution of joint video-action prediction.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

FutureNav: Unified World-Action Modeling for Vision-and-Language Navigation

Zhang, Lingfeng; Gong, Zeying; Hao, Xiaoshuai; Fu, Haoxiang; Zhang, Qiang; Zhou, Mingliang; Ye, Hangjun; Liang, Xiaojun; Liang, Junwei; Ding, Wenbo

FutureNav teaches a shared navigation VLM to predict actions and understand adjacent spatial transitions through three auxiliary tasks. Frozen VGGT features enrich its input and provide future-state targets. Default inference decodes actions without running the auxiliary heads. Simulated navigation and objective ablations provide quantitative evidence; physical transfer is illustrated qualitatively. Evidence: architecture, encoding, training-inference, main-results, objective-ablation and qualitative.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Pondering the Way: Spatial-perceiving World Action Model for Embodied Navigation

Chen, Hong; Liu, Daqi; Zhang, Zehan; Wang, Haiguang; Lu, Tianhao; Yan, Longfei; Sun, Haiyang; Li, Fangzhen; Xie, Hongwei; Wang, Bing; Chen, Guang; Ye, Hangjun; Tan, Yihua

SWAM turns start and goal images into a jointly generated RGB-D path and planar action sequence. A shared diffusion transformer, visual-guided action refinement, and endpoint regularization improve offline trajectory accuracy over candidate-ranking planners. Its evidence supports high-level planning, with mixed component gains and unresolved implementation details; it does not establish closed-loop robot execution.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

World Action Models Enable Continual Imitation Learning with Recurrent Generative Replays

Govind, Manish Kumar; Reilly, Dominick; Patel, Smit; Le, Hieu; Das, Srijan

REGEN reuses a world-action policy to generate rehearsal trajectories for old instructions while learning a new task. Current-task observations seed recurrent prediction; synthetic observation-action pairs then supplement new demonstrations. Retention improves substantially over sequential fine-tuning, but real replay generally remains stronger, and plausible imagined outcomes can disagree with executed actions.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

PAIWorld: A 3D-Consistent World Foundation Model for Robotic Manipulation

Yuhang Huang; Xuan Lv; Junyan Xu; Zhiyuan Yu; Jiazhao Zhang; Ruizhen Hu; Wancheng Feng; Shilong Zou; Hewen Xiao; Ziqiao Zhou; Kaiyun Huang; Zhiyu Peng; Juzhan Xu; Hang Zhao; Chenyang Zhu; Renjiao Yi; Yifei Huang; Douhui Wu; Yan Zhang; Kexu Cheng; Chunhe Song; Yunzhi Xue; Xiuhong Zhang; Leitao Guo; Yunji Chen; Bin Wu; Haibin Yu; Kai Xu

PAIWorld adds camera-aware communication and geometric representation distillation to a video diffusion transformer. It predicts future camera views from context plus text or supplied actions. Its strongest mechanism evidence is a four-setting ablation in which the combined additions reduce MEt3R more than either alone. The experiments establish generation improvements; claims about better robot planning and policies lack corresponding downstream evaluations in this PDF (e03, e04, e15, e17).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Cosmos 3: Omnimodal World Models for Physical AI

NVIDIA; Aditi; Niket Agarwal; Arslan Ali; Jon Allen; Martin Antolini; Adeline Aubame; Alisson Azzolini; Junjie Bai; Maciej Bala; Yogesh Balaji; Josh Bapst; Aarti Basant; Mukesh Beladiya; Mohammad Qazim Bhat; Zaid Pervaiz Bhat; Dan Blick; Vanni Brighella; Han Cai; Tiffany Cai; Eric Cameracci; Jiaxin Cao; Yulong Cao; Mark Carlson; Carlos Casanova; Ting-Yun Chang; Yan Chang; Yu-Wei Chao; Prithvijit Chattopadhyay; Roshan Chaudhari; Chieh-Yun Chen; Junyu Chen; Ke Chen; Qizhi Chen; Wenkai Chen; Xiaotong Chen; Yu Chen; An-Chieh Cheng; Click Cheng; Xiu Chia; Jeana Choi; Chaeyeon Chung; Wenyan Cong; Yin Cui; Magdalena Dadela; Nalin Dadhich; Wenliang Dai; Joyjit Daw; Alperen Degirmenci; Rodrigo Vieira Del Monte; Robert Denomme; Sameer Dharur; Marco Di Lucca; Ke Ding; Wenhao Ding; Yifan Ding; Yuzhu Dong; Nicole Drumheller; Yilun Du; Aigul Dzhumamuratova; Aleksandr Efitorov; Hamid Eghbalzadeh; Naomi Eigbe; Imad El Hanafi; Hassan Eslami; Benedikt Falk; Jiaojiao Fan; Jim Fan; Amol Fasale; Sergiy Fefilatyev; Liang Feng; Francesco Ferroni; Sanja Fidler; Xiao Fu; Vikram Fugro; Prashant Gaikwad; TJ Galda; Katelyn Gao; Yihuai Gao; Wenhang Ge; Sreyan Ghosh; Arushi Goel; Vivek Goel; Akash Gokul; Rama Govindaraju; Jinwei Gu; Miguel Guerrero; Elfie Guo; Aryaman Gupta; Siddharth Gururani; Hugo Hadfield; Song Han; Ankur Handa; Zekun Hao; Mohammad Harrim; Ali Hassani; Nathan Hayes-Roth; Yufan He; Chris Helvig; Cyrus Hogg; Madison Huang; Michael Huang; Sophia Huang; Yufan Huang; Jacob Huffman; DeLesley Hutchins; Suneel Indupuru; Boris Ivanovic; Arihant Jain; Joel Jang; Ryan Ji; Yanan Jian; Dongfu Jiang; Jingyi Jin; Atharva Joshi; Nikhilesh Joshi; Pranjali Joshi; Andy Ju; Jaehun Jung; Weiwei Kang; Scott Kassekert; Jan Kautz; Ashna Khetan; Julia Kiczka; Slawek Kierat; Gwanghyun Kim; Kuno Kim; Sunny Kim; Kezhi Kong; Xin Kong; Zhifeng Kong; Tomasz Kornuta; Egor Krivov; Hui Kuang; Saurav Kumar; Chia-Wen Kuo; George Kurian; Wojciech Kutak; JF Lafleche; Himangshu Lahkar; Omar Laymoun; Jayjun Lee; Sanggil Lee; Gabriele Leone; Boyi Li; Freya Li; Jiajun Li; Jinfeng Li; Ling Li; Pengcheng Li; Shangru Li; Tingle Li; Xiaolong Li; Xuan Li; Zhaoshuo Li; Zhiqi Li; Hao Liang; Maosheng Liao; Chen-Hsuan Lin; Tsung-Yi Lin; Ming-Yu Liu; Sifei Liu; Zihan Liu; Hai Loc Lu; Xiangyu Lu; Alice Luo; Ruipu Luo; Wenjie Luo; Jiangran Lyu; Martin Ding Ma; Nic Ma; Qianli Ma; Dawid Majchrowski; Louis Marcoux; Miguel Martin; Qing Miao; Ashkan Mirzaei; Shreyas Misra; Kaichun Mo; Durra Mohsin; Hyejin Moon; Pawel Morkisz; Saeid Motiian; Kirill Motkov; Seungjun Nah; Yashraj Narang; Deepak Narayanan; Thabang Ngazimbi; Julian Ouyang; Shubham Pachori; David Page; Yatian Pang; Sehwi Park; Mahesh Patekar; Mostofa Patwary; Marco Pavone; Trung Pham; Wei Ping; Soha Pouya; Shrimai Prabhumoye; Varun Praveen; Delin Qu; Hesam Rabeti; Morteza Ramezanali; Marilyn Reeb; Xuanchi Ren; Kristen Rumley; Wojciech Rymer; Jun Saito; Yeongho Seol; John Shao; Piyush Shekdar; Tianwei Shen; Humphrey Shi; Min Shi; Stella Shi; Kevin Shih; Mohammad Shoeybi; Mateusz Sieniawski; Shuran Song; Alexander Sotelo; Amir Sotoodeh; Sunil Srinivasa; Vignesh Srinivasakumar; Bartosz Stefaniak; Rahul Heinrich Steiger; Shangkun Sun; Jiaxiang Tang; Shitao Tang; Yangyang Tang; Yue Tang; Tolou Tavakkoli; Kayley Ting; Krzysztof Tomala; Wei-Cheng Tseng; Jibin Varghese; Sergei Vasilev; Thomas Volk; Raju Wagwani; Roger Waleffe; Andrew Z. Wang; Boxiang Wang; Haoxiang Wang; Qiao Wang; Shihao Wang; Shijie Wang; Ting-Chun Wang; Yan Wang; Yu Wang; Rohit Watve; David Wehr; Fangyin Wei; Xinshuo Weng; Jay Zhangjie Wu; Kedi Wu; Hongchi Xia; Summer Xiao; Tianjun Xiao; Kevin Xie; Daguang Xu; Jiashu Xu; Mengyao Xu; Ruqing Xu; Xingqian Xu; Yao Xu; Dinghao Yang; Dong Yang; Hans Yang; Xiaodong Yang; Xuning Yang; Yichu Yang; Yurong You; Zhiding Yu; Hao Yuan; Simon Yuen; Xiaohui Zeng; Pengcuo Zeren; Cindy Zha; Haotian Zhang; Jenny Zhang; Jing Zhang; Liangkai Zhang; Paris Zhang; Shun Zhang; Xuanmeng Zhang; Zhizheng Zhang; Ann Zhao; Yilin Zhao; Yuliya Zhautouskaya; Charles Zhou; Fengzhe Zhou; Shilin Zhu; Yuke Zhu; Dima Zhylko; Artur Zolkowski

Cosmos 3 integrates autoregressive understanding with diffusion generation of images, video, audio and actions. Its most relevant world-action result is a shared generator that can predict future observations and executable robot actions jointly, then specialize through post-training. Broad benchmark gains support useful transfer, while mixed ablations and limited consistency diagnostics leave the strength of causal action grounding unresolved.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Critique of Agent Model

Eric Xing; Mingkai Deng; Jinyu Hou

Critique of Agent Model proposes Goal-Identity-Configurator (GIC): an agent that maintains subgoals and an evolving self-model, chooses when to simulate, and learns from real and imagined experience. A separately trained world model supplies dynamics. This is a conceptual and theoretical architecture paper; its aircraft pilot is a motivating example, and measured prototype results are deferred to companion work. The practical question is whether reliable simulation and learned regulation can justify their additional machinery. [e-architecture, e-separation, e-status]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

dVLA-RL: Reinforcement Learning over Denoising Trajectories for Discrete Diffusion Vision-Language-Action Models

Wu, Yuhao; Liu, Yitian; Shen, Weijie; Han, Mishuo; Xu, Wenjie; Liang, Haotian; Liu, Zhongshan; Mao, Yinan; Xu, Lei; Guan, Xinping; Ying, Ru; Zheng, Ran; Sui, Wei; Yang, Xiaokang; Ding, Wenbo; Mu, Yao

dVLA-RL applies online PPO to MM-ACT's discrete action denoising, scoring newly unmasked tokens across the sampled path. Its Hybrid variant assigns denoising budgets by task. Large reported simulation gains are peak online training-rollout results, and the practical likelihood drops scheduler probabilities (e-surrogate, e-hybrid, e-evaluation).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

A Watermark for Vision-Language-Action and World Action Models

Liu, Yule; Liu, Shuai; Wei, Jiaheng; He, Xinlei

Keyed latent-provenance verification marks selected sampling seeds in an existing robot policy, then recovers key evidence from executed action channels using the owner's differentiable generator. It supports both direct VLA and future-scene-mediated WAM policies. Sixteen-rollout groups separate the carried key in the clean tests, but calibration drift, strong output processing and replacement by a distilled student limit the ownership claim [e-threat, e-injection, e-map, e-main, e-calibration, e-distillation].

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Attacking the Trusted Imagination: Oracle-Level Integrity Attacks on Imagine-then-Act World Models

Chen, Linghan; Ji, Kaiyan; Guo, Minyu

This integrity-attack study perturbs observations to damage imagined futures in existing world-action models. Corruption is easier than cross-scene steering, and a denoiser detects many tested attacks. LaDi-WM MPC suffers a substantial single-task simulation success drop, but the shared visual pathway prevents attributing that failure exclusively to imagination.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Geometric Action Model for Robot Policy Learning

Han, Jisang; Jeon, Seonghu; Jung, Jaewoo; Zurbrügg, René; An, Honggyu; Portela, Tifanny; Hutter, Marco; Pollefeys, Marc; Kim, Seungryong; Hong, Sunghwan

GAM turns a pretrained geometric transformer into a language-conditioned manipulation policy by inserting temporal prediction between its shallow and deep blocks. Future geometry and action tokens share the decoder. Camera robustness is its clearest reported gain; ablations, runtime settings and internal reporting discrepancies qualify the broader claims.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

MV-WAM: Manifold-Aware World Action Model with Value Augmentation

Chen, Jintao; Jia, Peidong; Wuwu, Qingpo; Liu, Jiaming; Du, Mengfei; Fan, Chun-Kai; Chi, Xiaowei; Chen, Hao; Bai, Chengyu; Qian, Zezhong; Wang, Hao; Cao, Jiajun; Mi, Weishi; Ju, Xiaozhu; Tang, Jian; Zhang, Shanghang

MV-WAM couples video, action and value generation through asymmetric shared attention. It trains video on velocity targets and actions on clean endpoints, then uses predicted progress to trigger rollback. Its strongest evidence is improved RoboTwin Random success; geometric necessity and reliable physical rollback remain less established (e-architecture, e-objectives, e-value, e-simulation, e-theory).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

World Action Models: A Survey

Shen, Qiuhong; Zhang, Shihua; Liao, Yue; Li, Qi; Tan, Zhenxiong; Wang, Shizun; Yan, Shuicheng; Wang, Xinchao

This survey organizes embodied models by how future prediction becomes useful to action. Its two complementary views separate the obligation to render a future from the representation, predictor, action interface and control schedule implementing it. The authors argue for selective imagination: retain information that constrains executable control while accounting for latency, memory and action-label cost. This is a conceptual synthesis of cited systems, not a newly trained policy or a controlled performance comparison (E02, E03, E20).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

MemoryWAM: Efficient World Action Modeling with Persistent Memory

Yang, Sizhe; Mu, Juncheng; Wei, Tianming; Lu, Chenhao; Li, Xiaofan; Xu, Linning; Xue, Zhengrong; Yuan, Zhecheng; Lin, Dahua; Pang, Jiangmiao; Xu, Huazhe

MemoryWAM compresses older visual history into per-frame gist tokens while retaining full initial and recent observations. A video/action mixture of transformers learns with both prediction targets, then executes action chunks without generating future video. Its strongest evidence concerns memory-dependent manipulation and the efficiency of its cache design.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

EventVLA: Event-Driven Visual Evidence Memory for Long-Horizon Vision-Language-Action Policies

Ganlin Yang; Zhangzheng Tu; Yuqiang Yang; Sitong Mao; Junyi Dong; Tianxing Chen; Jiaqi Peng; Jing Xiong; Jiafei Cao; Jifeng Dai; Wengang Zhou; Yao Mu; Tai Wang

EventVLA augments a VLA controller with sparse images of earlier events. A parallel head forecasts which upcoming action steps merit a memory write; the system saves the actual image when that step occurs. Its diagnostic RoboTwin-MeM result rises from 18.0% with fixed anchors to 75.2% with KEM. This supports memory-assisted robot control, with no future-image prediction objective [e-architecture, e-robomem].

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?

Zhang, Yuyang; Zhang, Wenyao; Qi, Zekun; Zhang, He; Lin, Haitao; Zhang, Jingbo; Mu, Yao; Yang, Xiaokang; Zeng, Wenjun; Jin, Xin

ImageWAM adapts pretrained image-editing models into robot policies: future endpoint supervision shapes internal visual features, and a separate flow-matching action expert reads their key/value caches. Deployment uses one editing forward pass without decoding the endpoint image. Reported robustness and efficiency gains depend strongly on the backbone and comparison protocol.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

WAM-RL: World-Action Model Reinforcement Learning with Reconstruction Rewards and Online Video SFT

Qian, Zezhong; Chi, Xiaowei; Qi, Yu; Li, Haozhan; Chen, Zhi Yang; Zhang, Shanghang

WAM-RL post-trains a video world model and a separate actor through environment interaction. Successful rollouts refine video prediction under a latent KL constraint; reconstruction rewards teach the actor to realize imagined futures. Reported success rises from 68% to 82% on LIBERO-Object and from 19% to 22% on RLBench Water Plants. These modestly documented benchmark results motivate coordinated adaptation, but do not establish a general long-horizon remedy. [e-framework, e-kl, e-reward, e-main]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Metis: A Generalizable and Efficient World-Action Model for Autonomous Driving and Urban Navigation

Li, Jingyu; Liu, Zhe; Hu, Dongnan; Wu, Junjie; Ma, Zipei; Wu, Wenxiao; Han, Chao; Hao, Zhihui; Liu, Zhikang; Zhan, Kun; Deng, Jiankang; Zhu, Xiatian; Zhang, Li

Metis trains separate video and action transformer experts together, then predicts trajectories without synthesizing future video. Its asymmetric attention supplies a training-time path from video supervision into the action expert. Driving and navigation results support this design, but latency comparisons and several implementation details require careful qualification.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies

Chen, Jialei; Wang, Kai; Chen, Kang; Chen, Shuaihang; Gao, Feng; Tang, Wenhao; Li, Zhiyuan; Liu, Weilin; Yao, Zhuyu; Li, Boxun; Xu, Yuanbo; Yu, Chao

LaWAM turns a latent-action model's forward decoder into an inference-time source of visual subgoals. A language-conditioned policy predicts a latent transition, LaWM expands it into DINO future features, and a separate action expert generates a robot action chunk. The main appeal is near-leading manipulation success with inexpensive future prediction; the evidence remains concentrated on stable-camera manipulation.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Retrieve, Don't Retrain: Extending Vision Language Action Models to New Tasks at Test Time

Park, Jeongeun; Park, Juhan; Kim, Taekyung; Choi, Sungjoon; Han, Dongyoon; Yun, Sangdoo

ReCAP freezes a retrieval-conditioned Cosmos Policy after paired cross-embodiment training. New pool trajectories supply task progression; a learned residual adapts their actions to the target robot, jointly with future-image prediction. Unseen-task execution improves on PushT, RoboTwin and a small real-robot study, but requires compatible action coordinates and curated trajectory data (e-adaptation, e-architecture, e-pusht-results, e-robotwin-results, e-real-results, e-limits).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

MemoryVAM: Integrating Memory into Video Action Model for Robot Manipulation

Yuxin Jiang; Chang Yu; Yunuo Chen; Xiang Feng; Yin Yang; Nishank Gite; Chenfanfu Jiang

MemoryVAM compresses episode history into tokens shared by a video predictor and downstream action decoder, with a learned completion gate. On LIBERO-Mem, UNet success rises from 5.0% to 42.5%; physical DiT rollouts also improve. The evidence supports history-conditioned predictive representations, while long-count reliability, scaling and exact reproduction remain unresolved (e02, e03, e10, e11, e15, e17).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

RepWAM: World Action Modeling with Representation Visual-Action Tokenizers

Wang, Junke; Zhang, Qihang; Yang, Shuai; Luo, Yiming; Shen, Yujun; Wu, Zuxuan; Jiang, Yu-Gang; Xu, Yinghao

RepWAM learns semantic visual tokens and compact transition tokens, jointly models their future evolution under language, then adapts to robot commands. Its central evidence is the tokenizer and training-stage ablations: better latent prediction accompanies better physical execution. The main simulation comparison remains below Lingbot-VA, and robot results use few trials.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

WAM-Nav: Asymmetric Latent World-Action Modeling for Unified Visual Navigation

Yang, Ning; Huang, Yan; Peng, Kaiwen; He, Ziheng; Wang, Kai; Miao, Cui; Lyu, Kailin; Li, Guo; Wang, Xiaofeng; Zhu, Zheng; Liu, Jing; Liu, Nianfeng

WAM-Nav jointly generates navigation trajectories and short-horizon visual latents using one shared Diffusion Transformer. Goal-conditioned visual and motion histories guide a 24-step action horizon with one future visual state. Reported zero-shot navigation improves over NavDP, but task balancing sacrifices specialized peaks, and physical failures expose camera visibility and body-clearance limits. Method, results and boundaries are traced in e02–e18.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models

Yu, Hanyang; Lin, Haitao; Zhang, Jingbo; Zhang, Wenyao; Gu, Chenghao; Li, Heng; Tan, Ping

MaskWAM couples RGB and mask future prediction with action denoising, using an optional first-frame target mask to resolve ambiguous instructions. Its central empirical distinction is between supplying a mask and training the policy to predict its future: the latter substantially improves mask-guided robot execution in the evaluated setup. Deployment uses partially denoised visual latents, while fully decoded future videos are offline illustrations (joint, inference, ablation).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

NavWAM: A Navigation World Action Model for Goal-Conditioned Visual Navigation

Azuma, Daichi; Miyanishi, Taiki; Sakamoto, Koya; Kurita, Shuhei; Zhu, Yaonan; Khrapchenkov, Petr; Kawanabe, Motoaki; Iwasawa, Yusuke; Matsuo, Yutaka

NavWAM adapts a video diffusion transformer into an image-goal navigation policy by jointly denoising actions, future views, state and goal progress. Its strongest mechanism evidence is a controlled future-image-loss ablation; its physical deployment reports 19/24 successes. The default policy avoids candidate-action search, but the study does not establish calibrated foresight or broad real-world robustness.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Diffusion Transformer World-Action Model for AV Scene Prediction

Sharifullin, Ruslan; Jiang, Benjamin; Chew, Kai Xi

This compact driving world model predicts future front-camera latents from a present frame and supplied ego-actions. Calibrated diffusion improves distributional appearance over deterministic regression, while coherent motion remains weak. A separate chained jump model recovers coarse motion but accumulates blur. These are offline prediction and controllability findings, with closed-loop driving left untested (e-architecture, e-realism, e-jump-result, e-limits).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

FAWAM: Force-Aware World Action Models for Closed-Loop Contact-Rich Manipulation

He, Haotian; Yan, Zeyu; Liu, Qipeng; Guo, Ning; Lian, Wenzhao

FAWAM predicts action chunks and wrist wrenches from a video-derived representation, then uses measured–predicted wrench discrepancies to gate faster residual corrections during real robot execution. Its strongest evidence is improved completion of four contact-rich tasks; the full system also requires additional human-intervention training data.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

EWAM: An Enhanced World Action Model for Closed-Loop Online Adaptation in Embodied Intelligence

Zhou, Xin; Miao, Cong

EWAM adds memory-conditioned generation, anomaly detection, discrete routing and action correction to a frozen Cosmos3-Nano--Policy-DROID policy. Its local advantage is execution efficiency: BananaInBowlTask time falls from 25.60 to 9.27 seconds while success stays at 100%. This adaptation uses accumulated simulation experience and lightweight updates, and increases inference latency. The evidence supports a bounded simulation deployment result, with unresolved figure and protocol inconsistencies. [e02, e10, e13, e15, e17]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

World Pilot: Steering Vision-Language-Action Models with World-Action Priors

Lin, Zefu; Cui, Rongxu; Xu, Junjia; Jin, Xiaojuan; Li, Wenling; Fan, Lue; Zhang, Zhaoxiang

World Pilot adds two frozen world-action-model priors to an ABot-M0 VLA: future-scene latents update visual-language hidden states, while one encoded trajectory token conditions the action generator. The VLA still predicts executable actions under expert supervision. Reported gains concentrate on OOD manipulation, including 84.7% LIBERO-Plus success. The design improves modular reuse of video-derived knowledge but requires an additional world-model pass at every decision step.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Making Foresight Actionable: Repurposing Representation Alignment in World Action Models

Qiu, Lu; Li, Yizhuo; Chen, Yi; Ge, Yuying; Ge, Yixiao; Liu, Xihui

AGRA improves a video-conditioned robot policy by aligning intermediate video features with frozen DINOv2 features during training. A multi-layer bridge preserves predictive information for action decoding. Physical Pick-and-Place success rises from 34% to 80%; controlled simulation gains are smaller. The evidence supports a useful interface regularizer, with limited task coverage and incomplete causal characterization.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Efficient-WAM: A 1B-Parameter World-Action Model with Low-Cost Future Imagination

Li, Jiajun; Guo, Tiecheng; Ye, Yifan; Zhang, Rongyu; Chi, Xiaowei; Sun, Qianpu; Li, Ying; Lou, Yunfan; Huang, Yan; Lu, Zhihe; Guo, Meng; Zhang, Shanghang

Efficient-WAM compresses a future-video expert and uses its coarse latent predictions to guide action generation. Its RT variant combines sparse future tokens with fewer video updates, trading some simulation success for faster policy calls. The evidence supports an efficiency–control tradeoff, without establishing that visual detail is universally dispensable.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Next Forcing: Causal World Modeling with Multi-Chunk Prediction

Xu, Gangwei; Zhang, Qihang; Zhou, Jiaming; Zhu, Xing; Shen, Yujun; Yang, Xin; Xu, Yinghao

Next Forcing trains a causal video/action model with auxiliary losses for three future video chunks. Chained predictors teach the video backbone across temporal horizons; actions still come from inverse dynamics inside a unified transformer. RoboTwin gains are largest early in high-frame-rate training. Auxiliary modules can be removed or one retained for two-chunk generation, with an accuracy tradeoff and unmeasured end-to-end speedup.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

HiMem-WAM: Hierarchical Memory-Gated World Action Models for Robotic Manipulation

Sun, Xiaoquan; Zhang, Ruijian; Cao, Chen; Sun, Yihan; Chen, Jiahui; Xu, Zetian; Chen, Bo; Chen, Haijier; Yang, Zhen; Zhu, Jiarun; Hong, Yijun; Xu, JingZhe; Pang, Jingrui; Yuan, Mingqi; Chen, Jiayu

HiMem-WAM trains motion latents from optical flow, groups them into skill latents, and uses learned skill transitions to supervise sparse external-memory writes. Its described deployed policy selects a skill, predicts a latent action chunk, and decodes robot controls using current observations and memory. Stage II ablations support improved manipulation robustness; comparisons on memory tasks do not isolate the gate's contribution. Backbone inconsistencies and missing implementation settings limit reproducibility.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

AHA-WAM:Asynchronous Horizon-Adaptive World-Action Modeling with Observation-Guided Context Routing

Cai, Jisong; Ling, Long; Chu, Shiwei; Liu, Zhongshan; Kang, Jiayue; Liang, Zhixuan; Xu, Wenjie; Mao, Yinan; Zhang, Weinan; Yang, Xiaokang; Ying, Ru; Zheng, Ran; Mu, Yao

AHA-WAM separates a video-trained world planner from a fast action executor. The planner exposes reusable latent context; current images route and edit that context before each action chunk. Future-video prediction supervises training but is removed at deployment. The evidence supports strong manipulation performance with substantial systems acceleration, subject to different pretraining and timing protocols across comparisons.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

MotionWAM: Towards Foundation World Action Models for Real-Time Humanoid Loco-Manipulation

Zheng, Jia; Ma, Teli; Fan, Yudong; Wang, Zifan; Yang, Shuo; Liang, Junwei

MotionWAM conditions a whole-body motion generator on features from one forward pass of a video world model. Its three-stage training recipe leads to 76.1% mean success on nine physical Unitree G1 tasks, versus 43.9% for the strongest tested baseline. The key tradeoff is access to video-trained representations without fully generating future frames; token-interface ambiguities and limited generalization tests constrain reproduction and interpretation.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

C3^3ache: Accelerating World Action Models with Cross Inference Chunk Cache

Zhao, Weisen; Nguyen, Lam; Lu, Zhicong; Shang, Yuzhang

C³ache accelerates Fast-WAM by carrying a DiT residual cache across successive action chunks at matching denoising steps. It changes inference without retraining. LIBERO reaches 2.51× total inference speedup at 97.10% success, but the same aggressive configuration sharply reduces RoboTwin success; refresh and late-step computation are therefore central to the result.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Dream-Tac: A Unified Tactile World Action Model for Contact-Rich Robot Manipulation

Lou, Yunfan; Ye, Yifan; Fu, Yankai; Cen, Jun; Chi, Xiaowei; Lyu, Yaoxu; Jia, Peidong; Han, Sirui; Lu, Zhihe; Zhang, Shanghang

Dream-Tac fine-tunes a shared video diffusion transformer to generate robot actions together with future camera and tactile latents. A deterministic tactile-change gate directs attention toward touch, while fused attention and diffusion caching reduce computational cost. Real robot experiments support the multimodal design, but do not isolate the benefit of predicting future touch from simply observing it.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

ImagineUAV: Aerial Vision-Language Navigation via World-Action Modeling and Kinodynamic Planning

Liu, Xuchen; Huang, Jiawei; Xia, Shihao; Liu, Bingxi; Cui, Jinqiang; Yang, Jiankun

ImagineUAV turns a language instruction and current camera view into an imagined flight video, recovers relative poses with a separate learned visual-odometry model, and optimizes the resulting route against local geometry and flight constraints. Its simulation advantage and physical flights support this cascade, while inference latency and yaw/orbit failures limit broader deployment claims.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

SIMPLE: Simulation-Based Policy Learning and Evaluation for Humanoid Loco-manipulation

Wei, Songlin; Ni, Zhenhao; Liu, Jie; Zhao, Zhenyu; Ye, Junjie; Jing, Hongyi; Xia, Junkai; Liu, Xiawei; Leong, Michael; Heng, Liang; Huang, Di; Wang, Yue

SIMPLE combines MuJoCo physics, Isaac Sim rendering, demonstration collection and a common humanoid policy interface. Its central contribution is an evaluation and training environment. Experiments show task-dependent policy rankings, benefits from visual diversity and more demonstrations, and simulation-trained Psi0 achieving 8/10 physical successes on each of two tasks under similar settings. These findings support a useful testbed, while leaving broad real-world ranking fidelity incompletely established.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Light-WAM: Efficient World Action Models with State-Fusion Action Decoding

Li, Ziang; Cheng, Dongzhou; Wang, Yibin; Wang, Shiyue; Xu, Xiaoyang; Weng, Lingxuan; Wang, Juan; Wang, Jiaqi

Light-WAM learns robot actions from a video backbone adapted with LoRA and sparse residual adapters. Future-video flow matching supplies training supervision at reduced latent resolution; inference pools current-observation states from several layers and directly regresses an action chunk. The reported efficiency gains accompany competitive LIBERO performance but lower RoboTwin and real-world success than key baselines.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Dreaming when Necessary: Advancing World Action Models with Adaptive Multi-Modal Reasoning

Tang, Yinzhou; Xu, Jingbo; Shang, Yu; Song, Zihao; Gao, Chen; Wu, Wei; Li, Yong

AdaWAM spends computation selectively: it updates subtask language when needed, predicts future visual latents when needed, and always generates an action chunk. Separate text, video and action modules are coordinated by a supervised router. The clearest evidence is improved long-horizon and compositional task success, with a qualified efficiency tradeoff rather than uniformly faster action generation.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis

Yang, Yi; Liu, Zhihong; Kou, Siqi; Chen, Yiyang; Hu, Yanzhe; Zhou, Jianbo; Zhao, Boyuan; Wei, Zhijie; Xia, Xiao; Li, Xueqi; Liu, Pengfei; Deng, Zhijie

WLA-0 combines language-based subtask memory with a compact dynamics representation shared by separate World and Action Experts. Future-image supervision improves the action representation during training; efficient deployment omits image generation. Its strongest mechanism-specific evidence concerns memory-dependent manipulation, while simulated video transfer succeeds unevenly and human-video transfer fails on average.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Flash-WAM: Modality-Aware Distillation for World Action Models

Akbari, Arman; Zhang, Ci; Akbari, Arash; Zhao, Lin; Chen, Yixiao; Chen, Weiwei; Zhang, Xuan; Yuan, Geng; Wang, Yanzhi

Flash-WAM accelerates LingBot-VA with different video and action consistency functions while retaining a shared transformer and video-conditioned inverse dynamics. It trades some task success for lower denoising cost. The headline 85.54% RoboTwin result uses one video and two action steps; the fastest one-plus-one configuration reaches 81.41%. Internal protocol and table conflicts qualify the comparisons (E02, E05, E07, E14, E15).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Unified Video-Action Joint Denoising for Dexterous Action and Data Generation

Wang, Dingrui; Wang, YuAn; Liu, Jinkun; Zhang, Yue; Piccinini, Mattia; Sun, Yu; Betz, Johannes

Donk adapts a Wan video denoiser to generate videos and bimanual MANO trajectories through one shared backbone. An observed image yields policy-style future prediction; omitting it yields paired synthetic experience after an auxiliary initializer supplies hand-camera geometry. Offline trajectory and video results support this dual interface, while robot execution and the downstream usefulness of generated data remain untested (e_modes, e_inference, e_action_results, e_text_results, e_gaps).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

GeoSem-WAM: Geometry- and Semantic-Aware World Action Models

Ma, Fulong; Peng, Daojie; Yue, Wenjun; Cao, Jiahang; Wang, Bintao; Zhang, Qiang; Ma, Jun

GeoSem-WAM adds future depth and semantic prediction losses to a video/action model, then removes those prediction heads for robot deployment. Its central evidence is improved manipulation success with structured training supervision. Simulation gains are small and some averages are internally inconsistent; physical background and height tests show larger gains. Architecture and result evidence: e3–e15.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

WALL-WM: Carving World Action Modeling at the Event Joints

Shalfun Li; Victor Yao; Charles Yang; Truth Qu; Regis Cheng; Ryan Yu; Howard Lu; Newton Von; Vincent Chen; Yohann Tang; Maeve Zhang; Ellie Ma; Gody Li; Sage Yang; Lorien Shu; J. W. Gao; Ethan Chen; Colin Ye; Yu Sun; Elise Mon; PS Zhang; Neo Li; Lily Li; James Wang; Ping Yang; Chris Pan; Lucy Liang; Hang Su; Roy Gan; Hao Wang; Qian Wang

WALL-WM trains on semantic events that align language, future video and robot trajectories. A Wan-derived multi-view video tower supplies intermediate features to a separate action denoiser; event-language execution and history-conditioned fixed chunks share this backbone. Its clearest reported advantage is Task Progress under cluttered, changing instructions, with much smaller gains on precision insertion. The evidence supports this combined system, with substantial data and evaluation qualifications.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

AttenA+: Rectifying Action Inequality in Robotic Foundation Models

Peng, Daojie; Ma, Fulong; Cao, Jiahang; Zhang, Qiang; Xie, Xupeng; Guo, Jian; Luo, Ping; Luo, Andrew F.; Zhou, Boyu; Ma, Jun

AttenA+ changes how robotic policies learn from demonstrations: slow action steps receive greater relative loss weight because the authors associate them with precision-sensitive manipulation. It adds no learned attention module. Reported simulation and Franka gains make the heuristic promising, but task-dependent ablations, checkpoint selection and deliberately speed-shaped demonstrations limit the generality of the evidence.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Beyond Task Success: Behavioral and Representational Diagnostics for WAM and VLA

Mai, Hung; Zhu, Bin; Do, Tuan

This paper evaluates how robot policies succeed, pairing executed-action and object diagnostics with sparse-autoencoder (SAE) feature analysis. Its strongest lesson is that success, smoothness, selectivity and inference cost measure different properties. WAM advantages vary by architecture and benchmark; predictive feature labels provide associative evidence, with unresolved definition and reporting inconsistencies limiting precise replication (e02, e11, e12, e14, e15, e18, e21, e24, e26).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

DriveWAM: Video Generative Priors Enable Scalable World-Action Modeling for Autonomous Driving

Shi, Chen; Xu, Jinrui; Shi, Shaoshuai; Sheng, Kehua; Zhang, Bo; Jiang, Li

DriveWAM turns a pretrained video diffusion transformer into a driving policy: it predicts a future video latent, then generates ego motion conditioned on that future. A frozen VLM supplies changing semantic guidance, while selective video/action caches bound historical memory. The evidence supports benchmark trajectory prediction and a useful cache tradeoff, with route-label, data-curation, and deployment limits.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

OASIS: Observation-Action Space Alignment via SE(3) Trajectory Prediction for Robotic Manipulation

Chen, Xinzhe; Ren, Sihua; Huang, Liqi; Sun, Haowen; Li, Mingyang; Chen, Xingyu; Liu, Zeyang; Lan, Xuguang

OASIS guides a visuomotor policy with hidden states supervised to predict future camera-frame end-effector poses. Metric-depth features support this intermediate, and a separate learned decoder produces executable action chunks. Results favor the design on LIBERO, CALVIN and two real robot platforms, while routing ablations provide the clearest mechanism evidence. Pose prediction remains a control aid, not a guarantee of successful execution (e-alignment, e-trajectory, e-libero, e-calvin, e-real-results, e-ablation).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Point Tracking Improves World Action Models

Guan, Jiarui; Zhao, Wenshuai; Pei, Yue; Chen, Ziliang; Solin, Arno; Kannala, Juho

JOPAT learns robot actions together with future image latents and 2D point tracks. A shared diffusion transformer lets explicit motion correspondences influence action sampling, while visibility supervision represents missing visual evidence. The paper reports strong LIBERO and SO-101 success rates, but precision insertion remains weak. Its central contribution is a richer predicted state for control, supported by modality ablations rather than by visual realism alone.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Key-Gram: Extensible World Knowledge for Embodied Manipulation

Fan, Jingjing; Li, Siyuan; Ren, Botao; Deng, Zhidong

Key-Gram augments π0 and π0.5 with instruction-indexed embedding memory. A parser extracts reusable phrases; hashed lookup retrieves priors that modulate visual hidden states before attention. The backbone supplies future-visual representations and context for a separate action expert. Reported gains are strongest under distribution expansion and compositional manipulation, but the evidence does not isolate future prediction or establish lifelong, update-free knowledge retention.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

How to Instruct Your Robot: Dense Language Annotations Power Robot Policy Learning

Kim, Bosung; Wang, Ruiyi; Acuna, David; Jung, Jaehun; Trevithick, Alexander; Cui, Brandon; Choi, Yejin; Ammanabrolu, Prithviraj

DeMiAn re-annotates existing demonstrations with four kinds of dense language, then trains an instructor to supply useful captions during control. Its clearest test result is a five-percentage-point RoboCasa VLA gain. Benefits depend on task, annotation and deployment protocol; all execution experiments are simulated.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

WorldVLN: Autoregressive World Action Model for Aerial Vision-Language Navigation

Zhao, Baining; Xu, Jiacheng; Feng, Weicheng; Zhang, Xin; Wang, Zhaolu; Wang, Haoyang; Ji, Shilong; Wang, Ziyou; Fang, Jianjie; Zheng, Zhiheng; Zhang, Weichen; Shang, Yu; Wu, Wei; Gao, Chen; Chen, Xinlei; Li, Yong

WorldVLN turns an autoregressive video prior into a closed-loop aerial navigation policy: predict a short latent future, decode waypoint actions, execute them, and condition the next prediction on actual observations. Supervised grounding and online Action-aware GRPO improve reported outdoor and indoor success. Evidence is strongest for short-range simulated control; physical transfer is qualitative and server-assisted.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

World Action Models: The Next Frontier in Embodied AI

Wang, Siyin; Shi, Junhao; Fu, Zhaoyang; He, Xinzhe; Liu, Feihong; Yang, Chenchen; Zhou, Yikang; Fei, Zhaoye; Gong, Jingjing; Fu, Jinlan; Shou, Mike Zheng; Huang, Xuanjing; Qiu, Xipeng; Jiang, Yu-Gang

This survey defines World Action Models through future-state prediction coupled to action generation, then organizes the literature by architecture, data, and evaluation. Its useful contribution is a vocabulary for tracing how imagined futures influence control. It does not establish a winning architecture: the authors identify missing matched comparisons and missing tests of whether executed actions actually follow predicted futures.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

The DAWN of World-Action Interactive Models

Lu, Hongbo; Yao, Liang; He, Chenghao; Wang, Haoyu; Gu, Xiang; Li, Xianfei; Liao, Wenlong; He, Tao; Peng, Pai

DAWN plans driving trajectories by repeatedly exchanging information between an action-conditioned latent World Predictor and a world-conditioned diffusion Action Denoiser. Its strongest evidence combines improved nuScenes trajectory metrics with NAVSIM coupling ablations. Shorter rollouts offer a measured quality–latency tradeoff, but the headline configuration uses a full 4-second rollout and NAVSIM v2 exposes weaker rule compliance.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models

Feng, Qiuxuan; Yu, Jiale; Liu, Jiaming; Jia, Yueru; Wu, Zhuangzhe; Chen, Hao; Qian, Zezhong; Gu, Shuo; Jia, Peng; Ma, Siwei; Zhang, Shanghang

HarmoWAM routes a shared video world model into two action experts: inverse-dynamics-style transit control and diffusion-based precise interaction. Real Franka experiments report strong generalization to changed backgrounds, positions and objects, but headline success rates average sub-stages and the implementation remains partly unspecified.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

When to Trust Imagination: Adaptive Action Execution for World Action Models

Wang, Rui; Zhang, Yue; Lin, Jiehong; Luo, Kuncheng; Wang, Jianan; Wang, Zhongrui; Qi, Xiaojuan

FFDC-WAM adds a learned execution verifier to Motus. It compares the current observation with cached predicted visual dynamics, planned actions, and instruction features, then continues or interrupts the rollout. RoboTwin results favor adaptive execution over short chunks, while physical experiments gain success at increased time and inference cost.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Geometry Guided Self-Consistency for Physical AI

Dai, Yinwei; Chen, Zhuofu; Yang, Lijie; Netravali, Ravi

KeyStone improves stochastic robot policies by sampling several action chunks from one shared context and executing a geometrically representative candidate. It adds no learned selector or policy training. Its strongest reported simulation gain is 13.3 percentage points, while its latency advantage depends on spare GPU capacity. Consensus can also reinforce an incorrect dominant mode. The supporting mechanism, results and boundaries are traced in e02–e13.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

NoiseGate: Learning Per-Latent Timestep Schedules as Information Gating in World Action Models

Huang, Wen; Sun, Haoran; Guo, Yongjian; Ma, Yunxuan; Li, Haoran; Long, Jing; Mo, Zhouying; Guan, Zhong; Guo, Yucheng; Di, Shuai; Xiong, Junwu

NoiseGate learns how quickly each imagined future video latent becomes clean while a joint video–action model generates robot actions. A reward-trained scheduler changes video timesteps while leaving the action clock fixed. The reported 50-task RoboTwin average rises from 92.58% to 94.28%; controlled ablations support scheduling benefits, but their printed baseline averages and the policy-density notation contain unresolved inconsistencies.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Is the Future Compatible? Diagnosing Dynamic Consistency in World Action Models

Ruan, Bo-Kai; Hsiao, Teng-Fang; Lo, Ling; Shuai, Hong-Han

This paper measures whether a world action model’s predicted future matches the observation produced by executing its actions, then uses agreement among sampled futures as a value-free selection proxy. Tests on Cosmos-Policy and LingBot-VA show modest simulated control gains, while background collapse demonstrates that predictable futures can still be unsuccessful. The contribution is a reliability diagnostic and inference procedure, rather than a newly trained WAM.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Learning Visual Feature-Based World Models via Residual Latent Action

Zhang, Xinyu; Xu, Zhengtong; Tao, Yutian; Wang, Yeping; She, Yu; Boularias, Abdeslam

RLA-WM compresses changes between DINO features into residual latent actions, then generates those latents from an observation and robot actions. It improves aggregate future-frame fidelity at far lower FLOPs than the video baseline, but costs more than direct DINO regression. Separate applications improve learning from mostly actionless demonstrations and explore policy optimization inside the learned model, with mixed robot-specific gains.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

OA-WAM: Object-Addressable World Action Model for Robust Robot Manipulation

Liu, Yushan; Sun, Peibo; Li, Shoujie; Xie, Yifan; Zhang, Lingfeng; Chao, Xintao; Dong, Shiyuan; Chen, Fang; Zhang, Xiao-Ping; Ding, Wenbo

OA-WAM gives a shared world/action transformer persistent object addresses alongside changing content. Its strongest evidence is improved simulated camera-shift robustness when address masking and residual resets are enabled. Future-slot prediction is auxiliary supervision, not an inference-time planner. Several implementation and aggregate-reporting inconsistencies limit reproducibility.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

CKT-WAM: Parameter-Efficient Context Knowledge Transfer Between World Action Models

Jiang, Yuhua; Guo, Yijun; Yang, Hongbing; Lei, Guojun; Chen, Nuo; Zhang, Yinuo; Yan, Shaoqiang; Lin, Bo; Gao, Feifei; Qi, Biqing

CKT-WAM trains a context-transfer module between frozen DreamZero-14B and Cosmos-Policy-2B backbones. One teacher observation pass supplies compressed, routed tokens to the student's text-conditioning pathway. The paper reports 86.1% LIBERO-Plus success and 83.3% average physical-task success. Its 187.4M trainable parameters equal 1.17% of the combined backbones, while teacher inference remains necessary. Main-text and appendix adapter specifications disagree (e-teacher, e-config, e-main-results, e-real, e-budget, e-layout-conflict).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

EA-WM: Event-Aware Generative World Model with Structured Kinematic-to-Visual Action Fields

Yang, Zhaoyang; Jin, Yurun; Qi, Lizhe; Huang, Cong; Chen, Kai

EA-WM turns robot kinematics into camera-aligned visual action fields and couples their denoising stream to a Wan2.2 video generator through event-supervised gates. It improves simulated rollout metrics while leaving precise action recovery and deployment unresolved.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

World Model for Robot Learning: A Comprehensive Survey

Bohan Hou; Gen Li; Jindou Jia; Tuo An; Xinying Guo; Sicong Leng; Haoran Geng; Yanjie Ze; Tatsuya Harada; Philip Torr; Oier Mees; Marc Pollefeys; Zhuang Liu; Jiajun Wu; Pieter Abbeel; Jitendra Malik; Yilun Du; Jianfei Yang

This survey organizes robotic world models by how prediction connects to action: as part of a policy, as a learned simulator, or as a generator of useful visual futures. Its central reading lesson is to track both the predictive representation and when it is used. Joint training, online imagination and real action execution are different claims. The benchmark compilation illustrates competitive methods across several architectures without establishing a universally superior design.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Being-H0.7: A Latent World-Action Model from Egocentric Videos

BeingBeyond Team (Hao Luo; Wanpeng Zhang; Yicheng Feng; Sipeng Zheng; Haiweng Xu; Chaoyi Xu; Ziheng Xi; Yuhui Fu; Zongqing Lu)

Being-H0.7 trains a robot policy to anticipate useful future information inside latent queries. A future-aware training branch aligns its hidden states with a deployable branch that sees only current context, avoiding visual rollout during control. Strong benchmark and real-robot results support the complete system, while missing component ablations leave the causal contribution of future alignment unresolved (E02–E09, E12, E15).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Motubrain: An Advanced World Action Model for Robot Control

Motubrain Team

Motubrain combines language, video and action streams in a unified generative robot policy. It transfers video priors into relative end-effector control, then uses asymmetric attention and asynchronous action chunks for deployment. Simulation, world-prediction and real-robot evaluations support different capabilities; the autoregressive attention specification remains internally inconsistent.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising

Jun Guo; Qiwei Li; Peiyan Li; Zilong Chen; Nan Sun; Yifei Su; Heyun Wang; Yuan Zhang; Xinghang Li; Huaping Liu

X-WAM adapts a pretrained video diffusion transformer to predict robot actions, future states and multi-view RGB-D observations together. A one-way depth branch adds geometric supervision, while asynchronous denoising releases actions before completing video generation. The paper reports strong simulated manipulation and reconstruction results plus a small physical earphone-packing evaluation. Its central tradeoff is useful spatial supervision without depth decoding during every action step; limited temporal context and delayed control remain unresolved.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Privileged Foresight Distillation: Zero-Cost Future Correction for World Action Models

Pengcheng Fang; Hongli Chen; Xiaohao Cai

Privileged Foresight Distillation (PFD) compares two attention masks on the same world-action backbone, then learns the future-enabled change in action velocity through an output adapter. Deployment uses only the current frame and corrected action denoising. The reported LIBERO mean improves by 1.15 percentage points over reproduced Fast-WAM, with a measured 5.15% cached-context latency overhead; the title's zero-cost wording does not mean zero additional computation.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

RL Token: Bootstrapping Online RL with Vision-Language-Action Models

Charles Xu; Jost Tobias Springenberg; Michael Equi; Ali Amin; Adnan Esmail; Sergey Levine; Liyiming Ke

RL Token (RLT) specializes a pretrained VLA through a compact representation and a small online actor–critic. A reconstruction-trained token supplies perceptual context, while sampled VLA actions guide useful refinements. Physical-robot experiments show faster critical phases and improved reliability, but rely on task demonstrations, human supervision, and carefully separated critical-phase and full-task evaluations. [e-token, e-policy, e-evaluation, e-results, e-human]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation

Feng Jiang; Yang Chen; Kyle Xu; Yuchen Liu; Haifeng Wang; Zhenhao Shen; Jasper Lu; Shengze Huang; Yuanfei Wang; Chen Xie; Ruihai Wu

RoboWM-Bench evaluates whether generated manipulation videos can be converted into robot actions that complete tasks in simulation. Separate human-hand retargeting and robot inverse-dynamics interfaces expose failures hidden by plausible imagery. Reliability tests support these interfaces on real demonstrations, but execution scores still depend on action extraction, reconstructed physics and task-specific checkers.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

π_0.7: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities

Physical Intelligence; Bo Ai; Ali Amin; Raichelle Aniceto; Ashwin Balakrishna; Greg Balke; Kevin Black; George Bokinsky; Shihao Cao; Thomas Charbonnier; Vedant Choudhary; Foster Collins; Ken Conley; Grace Connors; James Darpinian; Karan Dhabalia; Maitrayee Dhaka; Jared DiCarlo; Danny Driess; Michael Equi; Adnan Esmail; Yunhao Fang; Chelsea Finn; Catherine Glossop; Thomas Godden; Ivan Goryachev; Lachlan Groom; Haroun Habeeb; Hunter Hancock; Karol Hausman; Gashon Hussein; Victor Hwang; Brian Ichter; Connor Jacobsen; Szymon Jakubczak; Rowan Jen; Tim Jones; Gregg Kammerer; Ben Katz; Liyiming Ke; Mairbek Khadikov; Chandra Kuchi; Marinda Lamb; Devin LeBlanc; Brendon LeCount; Sergey Levine; Xinyu Li; Adrian Li-Bell; Vladislav Lialin; Zhonglin Liang; Wallace Lim; Yao Lu; Enyu Luo; Vishnu Mano; Nandan Marwaha; Aikys Mongush; Liam Murphy; Suraj Nair; Tyler Patterson; Karl Pertsch; Allen Z. Ren; Gavin Schelske; Charvi Sharma; Baifeng Shi; Lucy Xiaoyang Shi; Laura Smith; Jost Tobias Springenberg; Kyle Stachowicz; Will Stoeckle; Jiaming Tang; Jimmy Tanner; Shalom Tekeste; Marcel Torne; Kyle Vedder; Quan Vuong; Anna Walling; Haohuan Wang; Jason Wang; XuDong Wang; Chris Whalen; Samuel Whitmore; Blake Williams; Charles Xu; Sukwon Yoo; Lili Yu; Wuming Zhang; Zhuoyang Zhang; Ury Zhilinsky

π0.7 turns a generalist robot policy into a steerable controller by conditioning on detailed subtasks, desired episode quality and speed, and optional generated subgoal images. A separate world model supplies visual goals; the VLA predicts executable action chunks. The evidence supports strong performance on trained tasks and selected transfers, but novel long tasks can require coaching, and broad training exposure makes novelty difficult to certify. This report concerns the supplied April 24, 2026 v2.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

DexWorldModel: Causal Latent World Modeling towards Automated Learning of Embodied Tasks

Yueci Deng; Guiliang Liu; Kui Jia

DexWorldModel introduces CLWM, which predicts future DINOv3 features and then generates actions conditioned on that prediction. Shared transformer blocks, separate persistent and speculative memories, and asynchronous denoising connect world prediction to robot execution. EmbodiChain supplies synthetic adaptation data. Reported manipulation results are strong, but protocol omissions and a contradictory flow-time convention limit reproducibility.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

AIM: Intent-Aware Unified world action Modeling with Spatial Value Maps

Liaoyuan Fan; Zetian Xu; Chen Cao; Wenyao Zhang; Mingqi Yuan; Jiayu Chen

AIM jointly generates future RGB observations, spatial contact maps and continuous robot actions. A masked mixture-of-transformers architecture makes the predicted maps the action head's only route to future information. Supervised learning uses simulator-derived contact labels; post-training freezes world/value prediction and refines actions using map responses plus task rewards. The paper reports 94.0% Easy and 92.1% Hard success on RoboTwin 2.0, but the small post-training improvement, missing mechanism ablations and inconsistent schematic require careful interpretation.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis

Xiaolei Lang; Yang Wang; Yukun Zhou; Chaojun Ni; Kerui Li; Jiagang Zhu; Tianze Liu; Jiajun Lv; Xingxing Zuo; Yun Ye; Guan Huang; Xiaofeng Wang; Zheng Zhu

VAG synthesizes robot videos and action sequences from an initial image and instruction. A video diffusion transformer supplies pooled clean-latent predictions to an action U-Net during synchronized denoising. The evidence supports improved action prediction and simulation replay, plus a small real-robot policy-pretraining benefit. It does not establish guaranteed alignment or general closed-loop control.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Learning Vision-Language-Action World Models for Autonomous Driving

Guoqing Wang; Pin Tang; Xiangxuan Ren; Guodongfang Zhao; Bailan Feng; Chao Ma

VLA-World turns a predicted near-term ego motion into a generated camera view, then uses that imagined future as context for reasoning and a three-second waypoint plan. A Qwen2-VL-2B autoregressive model learns through visual pretraining, structured imitation and GRPO. Its nuScenes results improve planning and future-image FID, but establish offline prediction performance rather than closed-loop driving safety. The key research question is whether the generated future provides accurate decision-relevant evidence, rather than merely useful intermediate tokens.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

LMGenDrive: Bridging Multimodal Understanding and Generative World Modeling for End-to-End Driving

Hao Shao; Letian Wang; Yang Zhou; Yuxuan Hu; Zhuofan Zong; Steven L. Waslander; Wei Zhan; Hongsheng Li

LMGenDrive couples a language-based driving policy to a multi-view diffusion generator during training. Separate action and world queries connect instruction-grounded planning to future-video supervision. At deployment, online driving uses observed camera feedback and discards the generator; offline synthesis rolls generated views and predicted actions forward. The paper reports improved CARLA LangAuto driving scores, while its horizon experiment exposes accumulating video errors. This is evidence for a training-time generative contribution to simulated driving, with important implementation and evaluation gaps.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

ActiveGlasses: Learning Manipulation with Active Vision from Ego-centric Human Demonstration

Yanwen Zou; Chenyang Shi; Wenye Yu; Han Xue; Jun Lv; Ye Pan; Chuan Wen; Cewu Lu

ActiveGlasses transfers bare-hand demonstrations to robots by learning object motion and camera motion from a common 3D representation. Human stereo video supplies geometry and object-trajectory supervision; a robot arm then moves the same camera while another manipulates the object. Its contribution is an imitation-learning and data-collection system, with measured physical-task success rather than generated-video evaluation. The strongest controlled evidence favors active vision within the authors' policy; comparisons with π0.5 also change representation and demonstration collection [e-method, e-protocol, e-results].

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

INSPATIO-WORLD: A Real-Time 4D World Simulator via Spatiotemporal Autoregressive Modeling

InSpatio Team; Donghui Shen; Guofeng Zhang; Haomin Liu; Haoyu Ji; Hujun Bao; Hongjia Zhai; Jialin Liu; Jing Guo; Nan Wang; Siji Pan; Weihong Pan; Weijian Xie; Xianbin Liu; Xiaojun Xiang; Xiaoyu Zhang; Xinyu Chen; Yifu Wang; Yipeng Chen; Zhenzhou Fan; Zhewen Le; Zhichao Ye; Ziqiang Zhao

INSPATIO-WORLD turns a reference video into camera-controlled generated video. Its STAR generator combines a reference/history cache with depth-based reprojection; JDMD distills motion and appearance guidance into shared student weights. Benchmark gains support this generative simulator, while the small-model speed claim and larger-model quality results require separate interpretation.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Action Images: End-to-End Policy Learning via Multiview Video Generation

Haoyu Zhen; Zixian Gao; Qiao Sun; Yilin Zhao; Yuncong Yang; Yilun Du; Pengsheng Guo; Tsun-Hsuan Wang; Yi-Ling Qiao; Chuang Gan

Action Images makes robot control a multiview video-generation target: a shared backbone predicts future observations and RGB action heatmaps, which geometry converts into 7-DoF controls. Reported transfer is open-loop; an optional learned head improves in-domain success, and conflicting implementation specifications limit exact reproduction.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

JailWAM: Jailbreaking World Action Models in Robot Control

Hanqing Liu; Songping Wang; Jiahuan Long; Jiacheng Hou; Jialiang Sun; Chao Li; Yang Yang; Wei Peng; Xu Liu; Tingsong Jiang; Yao Mu; Wen Yao

JailWAM evaluates instruction-induced robot risk by rendering predicted actions as trajectory charts, screening them with a trained vision-language discriminator, and verifying selected candidates in closed-loop simulation. It reports 84.20% human-verified attack success on LingBot-VA, but its cheaper screening pipeline misses some unsafe executions. Success includes motion failure as well as catastrophic risk.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

DriveVA: Video Action Models are Zero-Shot Drivers

Mengmeng Liu; Diankun Zhang; Jiuming Liu; Jianfeng Cui; Hongwei Xie; Guang Chen; Hangjun Ye; Michael Ying Yang; Francesco Nex; Hao Cheng

DriveVA adapts a pretrained video generator to denoise future video and ego-trajectory tokens together. Its strongest evidence combines NAVSIM planning scores, direct transfer to nuScenes and CARLA, and ablations of video supervision. Jointly consistent outputs can still choose the wrong behavior; closed-loop evidence remains preliminary.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

SpatialVAM:Spatial-Aware Multi-View Video Diffusion as a Data-Efficient Robot Policy

Peiyan Li; Yixiang Chen; Yuan Xu; Jiabing Yang; Xiangnan Wu; Jun Guo; Nan Sun; Long Qian; Xinghang Li; Xin Xiao; Jing Liu; Nianfeng Liu; Tao Kong; Yan Huang; Liang Wang; Tieniu Tan

SpatialVAM turns a colored point cloud and robot position into aligned multiview RGB images and heatmaps, then adapts a pretrained video model to predict both future sequences. Geometric decoding recovers translation, while a separate latent decoder supplies rotation and gripper commands. The method performs strongly with few demonstrations, but depends on pretrained weights, calibrated depth sensing and expensive inference; visual plausibility alone does not establish safe execution.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

UniDriveVLA: Unifying Understanding, Perception, and Action Planning for Autonomous Driving

Yongkang Li; Lijun Zhou; Sixu Yan; Bencheng Liao; Tianyi Yan; Kaixin Xiong; Long Chen; Hongwei Xie; Bing Wang; Guang Chen; Hangjun Ye; Wenyu Liu; Haiyang Sun; Xinggang Wang

UniDriveVLA combines semantic understanding, sparse spatial perception and continuous trajectory generation in one Mixture-of-Transformers model. Separate expert parameters and directed attention aim to limit interference. Reported gains are strongest for selected planning and understanding comparisons; comfort, motion forecasting and preservation of foundation-model capability remain weaker.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

DriveDreamer-Policy: A Geometry-Grounded World-Action Model for Unified Generation and Planning

Yang Zhou; Xiaofeng Wang; Hao Shao; Letian Wang; Guosheng Zhao; Jiangnan Shao; Jiagang Zhu; Tingdong Yu; Zheng Zhu; Guan Huang; Steven L. Waslander

DriveDreamer-Policy trains a shared multimodal backbone with depth, video and trajectory generators. Ordered query embeddings transfer geometry and future-scene context into planning without requiring rendered depth or video at inference. Navsim scores and controlled modality ablations support useful joint supervision, while depth evaluation against a learned teacher and incomplete runtime details limit the conclusions.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Enhancing Policy Learning with World-Action Model

Yuci Han; Alper Yilmaz

WAM adds inverse action prediction between consecutive encoder embeddings to a DreamerV2 world model, then trains a separate diffusion policy on its frozen latent features. Table III supports 61.7% versus 45.8% average behavioral-cloning success on eight CALVIN tasks. The mechanism is plausible, but inconsistent headline numbers, baseline names and PPO summaries limit stronger conclusions.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

ManipArena: A Controlled Benchmark for Diagnosing Generalization in Real-Robot Manipulation

Yu Sun; Meng Cao; Yang Ping; Kaidong Zhang; Qingxuan Chen; Rongtao Xu; Liangwang Ruan; Xuecheng Chen; Dongxiu Liu; Yunxiao Yan; Zunnan Xu; Runze Xu; Charles Yang; Peilun Zhang; Xiaofan Li; Ruyi Gan; Liang Ma; Yuehao Yin; Jincheng Yu; Lufang Chen; Yuxin Liang; Peng Zhai; Hao Wang; Ivan Laptev; Ian Reid; Qian Wang; Xiaodan Liang

ManipArena evaluates manipulation policies through controlled physical robot trials, task schemas and subgoal scoring. Its strongest lesson is that training recipes and model provenance affect rankings alongside architecture. Language grounding and demonstration selection produce substantial reported gains, but small trial counts, restricted environments and internal reporting inconsistencies limit causal and generalization claims.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Uni-World VLA: Interleaved World Modeling and Planning for Autonomous Driving

Qiqi Liu; Huan Xu; Jingyu Li; Bin Sun; Zhihui Hao; Dangen She; Xiatian Zhu; Li Zhang

Uni-World VLA shares a multimodal backbone between future-frame generation and driving-trajectory prediction. It alternates imagined frames with waypoint queries and enriches historical RGB embeddings with estimated depth. NAVSIM results favor aligned 2-Hz generation, but the evidence concerns benchmark planning and generated imagery, not demonstrated physical deployment.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Vega: Learning to Drive with Natural Language Instructions

Sicheng Zuo; Yuxuan Li; Wenzhao Zheng; Zheng Zhu; Jie Zhou; Jiwen Lu

Vega learns instruction-conditioned driving by training trajectory denoising together with future-image denoising. Modality-specific transformers exchange information through global causal attention. InstructScene supplies automatically generated instructions describing recorded driving. The strongest reported NAVSIM v2 score uses best-of-six trajectory selection; NAVSIM v1 results are weaker than leading VLA baselines. Future-prediction ablations support visual supervision, while selected images illustrate instruction sensitivity without measuring general instruction-following reliability.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Persistent Robot World Models: Stabilizing Multi-Step Rollouts via Reinforcement Learning

Jai Bardhan; Patrik Drozdik; Josef Sivic; Vladimir Petrik

PersistWorld post-trains Ctrl-World on its own imperfect video histories. It branches comparable futures, scores them against recorded multi-view trajectories, and applies reward-weighted contrastive denoising. The evidence supports better visual persistence and modest downstream policy benefits, with unresolved implementation details and no explicit physical-consistency guarantee.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

DreamerAD: Efficient Reinforcement Learning via Latent World Model for Autonomous Driving

Pengxuan Yang; Yupeng Zheng; Deheng Qian; Zebin Xing; Qichao Zhang; Linbo Wang; Yichen Zhang; Shaoyu Guo; Zhongpu Xia; Qiang Chen; Junyu Han; Lingyun Xu; Yifeng Pan; Dongbin Zhao

DreamerAD makes Epona-based driving imagination cheaper, learns rewards directly from predicted latents, and uses those rewards to refine a trajectory policy with GRPO. Shortcut forcing enables one-step visual generation; vocabulary sampling limits exploration to structured trajectories. NAVSIM v2 EPDMS rises from 85.1 to 87.7, while ego progress falls. These are simulation planning results, with unresolved reward-implementation details.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Latent-WAM: Latent World Action Modeling for End-to-End Autonomous Driving

Linbo Wang; Yupeng Zheng; Qiang Chen; Shiwei Li; Yichen Zhang; Zebin Xing; Qichao Zhang; Xiang Li; Deheng Qian; Pengxuan Yang; Yihang Dong; Ce Hao; Xiaoqing Ye; Junyu han; Yifeng Pan; Dongbin Zhao

Latent-WAM trains a compact camera-based driving planner using geometric feature distillation and future latent-state prediction. Its deployment choice is to discard the dynamic predictor and decode trajectories from the current scene–ego representation. Reported gains concern NAVSIM planning scores and zero-shot HUGSIM simulation, with an HD-Score tie rather than an outright win (e03, e08, e10, e11).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Toward Physically Consistent Driving Video World Models under Challenging Trajectories

Jiawei Zhou; Zhenxin Zhu; Lingyi Du; Linye Lyu; Lijun Zhou; Zhanqian Wu; Hongcheng Luo; Zhuotao Tian; Bing Wang; Guang Chen; Hangjun Ye; Haiyang Sun; Yu Li

PhyGenesis first repairs potentially impossible driving trajectories, then renders their consequences as multi-view video. Counterfactual CARLA trajectories supervise a 6-DoF rectifier, while real and simulated clips train a layout-conditioned Wan2.1 generator. Its strongest gains concern challenging simulated inputs; the evidence establishes improved trajectory and video metrics, with substantial evaluation and reproducibility qualifications.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs

Haoran Yuan; Weigang Yi; Zhenyu Zhang; Wendi Chen; Yuchen Mo; Jiashi Yin; Xinzhuo Li; Xiangyu Zeng; Chuan Wen; Cewu Lu; Katherine Driggs-Campbell; Ismini Lourentzou

VTAM adapts a video world model to predict camera and tactile streams, then trains a conditional action expert with an auxiliary deformation-derived force target. It reports large gains in three real robot contact tasks. The strongest evidence is task success and a chip ablation; calibrated force accuracy, broad generalization and the claimed gradient mechanism remain unestablished. Method details below preserve inconsistencies between the formulation and implementation appendix.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Do World Action Models Generalize Better than VLAs? A Robustness Study

Zhanguang Zhang; Zhiyuan Li; Behnam Rahmati; Rui Heng Yang; Yintao Ma; Amir Rasouli; Sajjad Pakdamansavoji; Yangzheng Wu; Lingfeng Zhang; Tongtong Cao; Feng Wen; Xinyu Wang; Xingyue Quan; Yingxue Zhang

This study extends RoboTwin 2.0 with controlled perturbations and compares video-based world action models with VLAs on two simulated manipulation benchmarks. WAMs have strong visual-perturbation results, but π0.5 leads the reported LIBERO-Plus total. Unequal training histories, slow inference and an unresolved RoboTwin aggregation inconsistency limit broad claims about architectural superiority.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

X-World: Controllable Ego-Centric Multi-Camera World Models for Scalable End-to-End Driving

Chaoda Zheng; Sean Li; Jinhao Deng; Zhennan Wang; Shijia Chen; Liqiang Xiao; Ziheng Chi; Hongbin Lin; Kangjie Chen; Boyang Wang; Yu Zhang; Xianming Liu

X-World turns proposed driving actions and camera history into future surround-view video. A WAN-based latent generator combines camera/time attention with distinct control interfaces, then undergoes causal self-forcing training for streaming rollouts. Illustrative driving scenarios include a 24-second sequence; quantitative simulator fidelity and policy improvement remain unestablished (e-interface, e-architecture, e-stage2, e-long, e-evaluation).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

VAMPO: Policy Optimization for Improving Visual Dynamics in Video Action Models

Zirui Ge; Pengxiang Ding; Baohua Yin; Qishen Wang; Zhiyong Xie; Yemin Wang; Jinbo Wang; Hengtao Li; Runze Suo; Wenxuan Song; Han Zhao; Shangke Lyu; Zhaoxin Fan; Haoang Li; Ran Cheng; Cheng Chi; Huibin Ge; Yaozhi Luo; Donglin Wang

VAMPO post-trains a video predictor with rewards for agreement with expert future latents, then trains a separate action policy on its predictive features. First-step stochastic denoising concentrates optimization where the action policy obtains its representation. The supplied v1 reports stronger simulated manipulation and higher real-robot task scores, but latent agreement does not directly measure contact accuracy or guarantee control success.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

EVA: Aligning Video World Models with Executable Robot Actions via Inverse Dynamics Rewards

Ruixiang Wang; Qingming Liu; Yueci Deng; Guiliang Liu; Zhen Liu; Kui Jia

Executable Video Alignment (EVA) post-trains a video planner using a frozen inverse dynamics model to penalize rough or limit-violating decoded actions. Reported robot success improves, but smoothness remains an imperfect proxy for completing a task (e-overview, e-reward, e-real-results, e-hacking).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

GigaWorld-Policy: An Efficient Action-Centered World–Action Model

Angen Ye; Boyuan Wang; Chaojun Ni; Guan Huang; Guosheng Zhao; Hao Li; Hengtao Li; Jie Li; Jindi Lv; Jingyu Liu; Min Cao; Peng Li; Qiuping Deng; Wenjun Mei; Xiaofeng Wang; Xinze Chen; Xinyu Zhou; Yang Wang; Yifan Chang; Yifan Li; Yukun Zhou; Yun Ye; Zhichao Liu; Zheng Zhu

GigaWorld-Policy turns a video diffusion Transformer into a robot policy that learns actions alongside action-conditioned future visuals, then omits future-video generation during control. A shared causal mask makes this separation architectural. Physical experiments favor its speed–score tradeoff, although graded evaluation and inconsistent source numbers require care.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

DreamPlan: Efficient Reinforcement Fine-Tuning of Vision-Language Planners via Video World Models

Emily Yue-Ting Jia; Weiduo Yuan; Tianheng Shi; Vitor Guizilini; Jiageng Mao; Yue Wang

DreamPlan adapts a Qwen3-VL-8B manipulation planner using preferences derived from an action-conditioned video world model. Exploratory robot interactions train the dynamics predictor; predicted outcomes then support offline action ranking and planner fine-tuning. Deployment uses the adapted planner directly. On three physical deformable-object tasks, the reported mean score rises from 0.33 to 0.60 relative to the same zero-shot backbone, under a single-action evaluation that awards partial credit. The results support task-specific adaptation, while leaving verifier calibration and broader generalization unresolved.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Fast-WAM: Do World Action Models Need Test-time Future Imagination?

Tianyuan Yuan; Zibin Dong; Yicheng Liu; Hang Zhao

Fast-WAM trains a video transformer and an action expert together, then uses the video transformer only to encode the current observation during deployment. Action generation still uses iterative denoising. Controlled comparisons associate larger performance gains with the video training objective than with explicit test-time future generation, within the evaluated manipulation tasks (e-architecture, e-variants, e-robotwin, e-libero, e-real-results).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

S-VAM: Shortcut Video-Action Model by Self-Distilling Geometric and Semantic Foresight

Haodong Yan; Zhide Zhong; Jiaguan Zhu; Junjie He; Weilin Yuan; Wenxuan Song; Xin Gong; Yingjie Cai; Guanyi Zhao; Xu Yan; Bingbing Liu; Ying-Cong Chen; Haoang Li

S-VAM turns first-step video-diffusion features into predicted geometric and semantic representations, then conditions a separate diffusion policy on that foresight. Generated videos supervise the shortcut during training. Reported manipulation gains come with higher latency than VPP and do not establish equally fast visual feedback (e-shortcut, e-action, e-latency).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Bridging Scene Generation and Planning: Driving with World Model via Unifying Vision and Motion Representation

Xingtai Gui; Meijie Zhang; Tianyi Yan; Wencheng Han; Jiahao Gong; Feiyang Tan; Cheng-zhong Xu; Jianbing Shen

WorldDrive first learns trajectory-conditioned video dynamics, then transfers frozen vision and motion encoders into a trajectory planner. A separately trained Future-aware Rewarder approximates the world model's future latents to rank candidates without diffusion sampling at planning time. Its strongest evidence combines representation and rewarder ablations with single-camera NAVSIM results; generated counterfactual videos and oracle candidate scores have narrower evidential roles.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning

Lorenzo Mur-Labadia; Matthew Muckley; Amir Bar; Mido Assran; Koustuv Sinha; Mike Rabbat; Yann LeCun; Nicolas Ballas; Adrien Bardes

V-JEPA 2.1 learns image/video patch features by predicting both hidden and visible representations at multiple encoder depths. This substantially improves dense vision while largely preserving global recognition; separate task heads and world models turn the features into anticipation and planning systems.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels

Lucas Maes; Quentin Le Lidec; Damien Scieur; Yann LeCun; Randall Balestriero

LeWorldModel (LeWM) learns a compact, action-conditioned latent world model directly from offline image–action trajectories. Next-embedding prediction supplies dynamics supervision; SIGReg discourages representation collapse by matching random latent projections to a Gaussian target. A separate CEM solver chooses actions by predicting their consequences and matching a goal embedding. Strong Push-T performance and fast planning coexist with weaker results on Two-Room and OGBench-Cube than selected baselines. Physical probes and perturbation tests diagnose representations, rather than establish general physical reasoning.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

DiT4DiT: Jointly Modeling Video Dynamics and Actions for Generalizable Robot Control

Teli Ma; Jia Zheng; Zifan Wang; Chunli Jiang; Andy Cui; Junwei Liang; Shuo Yang

DiT4DiT conditions a separate action diffusion transformer on internal features of a video diffusion transformer and trains both with flow matching. One video-feature pass supports action inference without decoding future frames. Strong simulation and Unitree G1 results accompany slower deployment and unresolved timestep/configuration details (e04, e06, e07, e08, e11, e12, e13, e16, e17, e19).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

World2Act: Latent Action Post-Training from World Model Dynamics

An Dinh Vuong; Tuan Van Vo; Abdullah Sohail; Haoran Ding; Liang Ma; Xiaodan Liang; Anqing Duan; Ivan Laptev; Ian Reid

World2Act transfers an instruction-conditioned video world model's dynamics into a frozen GR00T-N1.6 policy through a learned video–action bridge and residual action-latent controller. Expert demonstrations first teach the bridge to align video and action chunks while reconstructing actions. Imagined latents then supervise residual corrections without IDM pseudo-actions or simulator rewards. Reported success improves on three simulation benchmarks and three physical tasks, but temporal correspondence remains necessary and plausible imagination can still fail at contact. The method and results below are from the supplied v2 artifact.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

SAMoE-VLA: A Scene Adaptive Mixture-of-Experts Vision-Language-Action Model for Autonomous Driving

Zihan You; Hongwei Liu; Chenxu Dang; Zhe Wang; Sining Ang; Aoqi Wang; Yan Wang

SAMoE-VLA conditions a flow-matching driving planner on world-language features and uses BEV scene context to merge expert parameters. Its strongest evidence is improved trajectory accuracy and LangAuto simulator driving scores; world rendering is pretraining-only, and several implementation and result descriptions conflict internally.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

MEM: Multi-Scale Embodied Memory for Vision Language Action Models

Marcel Torne; Karl Pertsch; Homer Walke; Kyle Vedder; Suraj Nair; Brian Ichter; Allen Z. Ren; Haohuan Wang; Jiaming Tang; Kyle Stachowicz; Karan Dhabalia; Michael Equi; Quan Vuong; Jost Tobias Springenberg; Sergey Levine; Chelsea Finn; Danny Driess

MEM adds two forms of history to π0.6: a recurrent language summary for semantic task progress and a video encoder for recent manipulation context. A high-level policy updates the summary and selects a subtask; a low-level policy produces robot actions. Physical-robot experiments support complementary memory benefits and correction-conditioned adaptation, while leaving training details and parts of the attention specification unresolved. [e-factorization, e-long-ablation, e-adaptation, e-attention-appendix]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Unifying Language-Action Understanding and Generation for Autonomous Driving

Xinyang Wang; Qian Liu; Wenjie Ding; Zhao Yang; Wei Li; Chang Liu; Bailin Li; Kun Zhan; Xianpeng Lang; Wei Chen

LinkVLA combines a shared language/action vocabulary, trajectory-to-language supervision and endpoint-conditioned parallel refinement. It improves reported CARLA driving and instruction-following scores, while its 48 ms trajectory-generation timing excludes textual reasoning. It predicts ego plans without learning future scene generation (e03, e06, e07, e10–e13).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Self-Correcting VLA: Online Action Refinement via Sparse World Imagination

Chenyv Liu; Wentao Tan; Lei Zhu; Fengling Li; Jingjing Li; Guoli Yang; Heng Tao Shen

SC-VLA adds progress and end-effector-change predictions to a GR00T N1.5-based flow policy, then freezes it and trains a SAC residual controller. Predicted motion supplies a directional reward whose influence decreases with predicted progress. Simulation supports both stages; physical ARX5 trials test only the predictive base policy. The evidence supports improved executed manipulation under the reported protocols, with unresolved reward and evaluation details.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data

Ruijie Zheng; Dantong Niu; Yuqi Xie; Jing Wang; Mengda Xu; Yunfan Jiang; Fernando Castañeda; Fengyuan Hu; You Liang Tan; Letian Fu; Trevor Darrell; Furong Huang; Yuke Zhu; Danfei Xu; Linxi Fan

EgoScale learns dexterous robot policies from action-labeled egocentric human video, then aligns them with robot sensing and control using a small human–robot play dataset. A vision-language backbone conditions a flow-based action expert predicting wrist motion and hand joints. Real-robot success improves after both human pretraining and aligned mid-training; transfer still requires robot demonstrations and embodiment-specific adaptation.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Learning to unfold cloth: Scaling up world models to deformable object manipulation

Jack Rome; Stephen James; Subramanian Ramamoorthy

This DreamerV2 adaptation learns to unfold cloth already suspended from one corner. Depth-derived surface normals, demonstration-seeded replay and temporally consistent image augmentation improve several simulated garment results. A plane-cloth policy transfers zero-shot to physical trials with 74% average success; this is not a universal garment policy. The method and deployment evidence are detailed in e02-task, e05-model, e07-replay and e13-physical.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Factored Latent Action World Models

Zizhao Wang; Chang Shi; Jiaheng Hu; Kevin Rohling; Roberto Martín-Martín; Amy Zhang; Peter Stone

FLAM learns a video world model whose slots each carry a latent action while sharing interaction-aware dynamics networks. Prediction-trained factorization improves rollouts supplied with future-inferred actions and can supply pseudo action labels for behavior cloning. Its strongest evidence concerns multi-entity video modeling; downstream control is evaluated separately in Procgen.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

WoVR: World Models as Reliable Simulators for Post-Training VLA Policies with RL

Zhennan Jiang; Shangqing Zhou; Yutong Jiang; Zefang Huang; Mingjie Wei; Yuhui Chen; Tianxing Zhou; Zhen Guo; Hao Lin; Quanlu Zhang; Yu Wang; Haoran Li; Chao Yu; Dongbin Zhao

WoVR post-trains a VLA policy inside an action-conditioned video simulator. Anchored visual memory stabilizes predictions, keyframe initialization shortens imagined prefixes, and low-frequency simulator refinement addresses policy drift. LIBERO and physical execution improve, but physical experiments use a different data protocol without PACE. Simulator reliability remains empirical.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion

Jiangran Lyu; Kai Liu; Xuheng Zhang; Haoran Liao; Yusen Feng; Wenxuan Zhu; Tingrui Shen; Jiayi Chen; Jiazhao Zhang; Yifei Dong; Wenbo Cui; Senmao Qi; Shuo Wang; Yixin Zheng; Mi Yan; Xuesong Shi; Haoran Li; Dongbin Zhao; Ming-Yu Liu; Zhizheng Zhang; Li Yi; Yizhou Wang; He Wang

LDA-1B learns policies and latent dynamics from heterogeneous embodied data by selecting supervision according to data quality. A shared multimodal diffusion transformer predicts actions or DINO future features under four task conditions. Simulation and physical manipulation results favor the system, but unequal fine-tuning data, partial-credit metrics and incomplete deployment details limit causal and reproducibility claims.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

GigaBrain-0.5M*: a VLA That Learns From World Model-Based Reinforcement Learning

GigaBrain Team; Boyuan Wang; Bohan Li; Chaojun Ni; Guan Huang; Guosheng Zhao; Hao Li; Jie Li; Jindi Lv; Jingyu Liu; Lv Feng; Mingming Yu; Peng Li; Qiuping Deng; Tianze Liu; Xinyu Zhou; Xinze Chen; Xiaofeng Wang; Yang Wang; Yifan Li; Yifei Nie; Yilong Li; Yukun Zhou; Yun Ye; Zhichao Liu; Zheng Zhu

GigaBrain-0.5M* adds a separate video world model to a robot VLA: predicted future latents supply scene context, while predicted values produce binary improvement labels for policy training. RAMP alternates this conditioning with physical rollouts and human corrections. Reported manipulation gains are promising, but incomplete protocols and chart–prose discrepancies limit precise interpretation.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

RISE: Self-Improving Robot Policy with Compositional World Model

Jiazhi Yang; Kunyang Lin; Jinwei Li; Wencong Zhang; Tianwei Lin; Longyan Wu; Zhizhong Su; Hao Zhao; Ya-Qin Zhang; Li Chen; Ping Luo; Xiangyu Yue; Hongyang Li

RISE improves a π0.5 robot policy using short imagined interactions. A separate action-conditioned video model predicts multiview observations; a learned value model scores their progress, producing advantage labels for policy training. Real-world tests cover moving-brick sorting, backpack packing and box closing. The world model is absent from deployment, but substantial offline experience and training compute remain necessary.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

VideoWorld 2: Learning Transferable Knowledge from Real-world Videos

Zhongwei Ren; Yunchao Wei; Xiao Yu; Guixun Luo; Yao Zhao; Bingyi Kang; Jiashi Feng; Xiaojie Jin

VideoWorld 2 learns discrete visual-dynamics codes with a pretrained diffusion appearance prior, then trains an autoregressive transformer to predict those codes. It improves generated long-horizon craft sequences and transfers latent pretraining to a separately action-supervised CALVIN policy. Its evidence supports benchmark-specific transfer; generated craft success, simulator control and complete appearance disentanglement remain distinct claims.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model

Jingwen Sun; Wenyao Zhang; Zekun Qi; Shaojie Ren; Zezhi Liu; Hanxin Zhu; Guangzhong Sun; Xin Jin; Zhibo Chen

VLA-JEPA trains a vision-language policy to produce latent actions that help a separate world model predict frozen video features, then learns continuous robot actions with flow matching. Its clearest gain is robustness to LIBERO-Plus perturbations; human-video benefits are weaker or reversed in other settings. The world model supplies training supervision, while the described deployment path generates actions without explicit future-state planning.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

WorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World Models

Yu Shang; Zhuohang Li; Yiding Ma; Weikang Su; Xin Jin; Ziyou Wang; Lei Jin; Xin Zhang; Yinzhou Tang; Haisheng Su; Chen Gao; Wei Wu; Xihui Liu; Dhruv Shah; Zhaoxiang Zhang; Zhibo Chen; Jun Zhu; Yonghong Tian; Tat-Seng Chua; Wenwu Zhu; Yong Li

WorldArena evaluates whether generated robot videos are useful for learning and control. It pairs sixteen video metrics with three functional tests: synthetic-data policy training, world-model policy evaluation, and action execution through an inverse dynamics model. Its EWMScore summarizes video quality, while separate RoboTwin experiments reveal weaker alignment with downstream performance. This is a benchmark contribution, not a newly unified robot policy.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

World-VLA-Loop: Closed-Loop Learning of Video World Model and VLA Policy

Xiaokang Liu; Zechen Bai; Hai Ci; Kevin Yuchen Ma; Mike Zheng Shou

World-VLA-Loop trains an action-conditioned video simulator to predict both future images and success rewards, then uses that simulator to improve a separate OpenVLA-OFT policy. Success and near-success demonstrations teach the simulator about subtle failures; fresh policy executions expand that data between training rounds. Two physical tasks improve, but simulator success remains distinct from execution success, and long rollouts degrade (e-method, e-loop, e-policy, e-iteration, e-horizon).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Visuo-Tactile World Models

Carolina Higuera; Sergio Arnaud; Byron Boots; Mustafa Mukadam; Francois Robert Hogan; Franziska Meier

VT-WM predicts future visual and fingertip-touch latents under candidate robot actions, then supplies those predictions to a separate goal-image CEM planner. Contact sensing improves rollout consistency and reported real-robot planning on familiar contact-rich tasks. The evidence supports this restricted setting; it does not establish novel-object generalization or high-frequency feedback control.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

HumanX: Toward Agile and Generalizable Humanoid Interaction Skills from Human Videos

Yinhuai Wang; Qihan Zhao; Yuen Fui Lau; Runyi Yu; Hok Wai Tsui; Qifeng Chen; Jingbo Wang; Jiangmiao Pang; Ping Tan

HumanX turns monocular demonstrations into humanoid control policies through XGen, which synthesizes interaction references using contact anchors and physics, and XMimic, which trains imitation policies with privileged teachers and deployable students. It supports proprioception-only basketball behaviors and externally tracked interactions. Its strongest evidence concerns simulated generalization within specified augmented distributions and physical Unitree G1 trials under distinct perception protocols (e02, e04, e07, e12, e13, e15, e16).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Causal World Modeling for Robot Control

Lin Li; Qihang Zhang; Yiming Luo; Shuai Yang; Ruilin Wang; Fei Han; Mingrui Yu; Zelin Gao; Nan Xue; Xing Zhu; Yujun Shen; Yinghao Xu

LingBot-VA turns a pretrained video generator into a robot policy: predict a short visual future, decode actions from that future and observation/action history, execute, and incorporate physical feedback. A unified Mixture-of-Transformers shares attention between video and action streams. Strong reported simulation results coexist with unresolved scoring and implementation inconsistencies.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning

Moo Jin Kim; Yihuai Gao; Tsung-Yi Lin; Yen-Chen Lin; Yunhao Ge; Grace Lam; Percy Liang; Shuran Song; Ming-Yu Liu; Chelsea Finn; Jinwei Gu

Cosmos Policy adapts Cosmos-Predict2-2B to generate robot action chunks, future observations and values in one latent diffusion sequence. Direct control benefits from joint supervision; optional planning adds rollout refinement and a separate planning checkpoint. Experiments support strong manipulation performance, while slow search and uneven OOD performance limit broader conclusions.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

UniDrive-WM: Unified Understanding, Planning and Generation World Model for Autonomous Driving

Zhexiao Xiong; Xin Ye; Burhan Yaman; Sheng Cheng; Yiren Lu; Jingru Luo; Nathan Jacobs; Liu Ren

UniDrive-WM couples a shared driving VLM to a waypoint planner and a future front-view image generator. Planning tokens precede visual prediction, allowing image supervision to improve shared representations. Discrete AR is faster; continuous AR+Diffusion yields lower FID. Bench2Drive closed-loop gains over ORION are modest, and do not establish real-vehicle safety or an inference-time imagination-and-replanning loop.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Wow, wo, val! A Comprehensive Embodied World Model Evaluation Turing Test

Chun-Kai Fan; Xiaowei Chi; Xiaozhu Ju; Hao Li; Yong Bao; Yu-Kai Wang; Lizhang Chen; Zhiyuan Jiang; Kuangzhi Ge; Ying Li; Weishi Mi; Qingpo Wuwu; Peidong Jia; Yulin Luo; Kevin Zhang; Zhiyuan Qin; Yong Dai; Sirui Han; Yike Guo; Shanghang Zhang; Jian Tang

WoW-World-Eval tests whether instruction-conditioned robot videos are visually convincing, task-correct, physically plausible and usable for action extraction. Its 609-sample benchmark combines automated metrics, human judgments and a separate GC-IDM robot-execution test. Hailuo leads the reported aggregate video score, while WoW-wan leads physical execution. The central lesson is that a plausible imagined manipulation and an executable one are distinct outcomes; calibration and evaluator dependence constrain how broadly the scores can be interpreted.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos

Chubin Zhang; Jianan Wang; Zifeng Gao; Yue Su; Tianru Dai; Cai Zhou; Jiwen Lu; Yansong Tang

CLAP turns human visual transitions into robot-grounded action tokens, then trains a language-conditioned policy on robot demonstrations and pseudo-labeled videos. Its optional rectified-flow controller replaces autoregressive decoding with faster continuous action chunks. The clearest transfer evidence concerns object selection within familiar tasks: human videos improve OOD bouquet success, but placement failures remain. Future-feature reconstruction supports representation training; this is not an inference-time world simulator.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation

Wenlong Huang; Yu-Wei Chao; Arsalan Mousavian; Ming-Yu Liu; Dieter Fox; Kaichun Mo; Li Fei-Fei

PointWorld forecasts how observed scene points move under a proposed robot trajectory. It places RGB-D scene geometry and forward-kinematic gripper motion in one point-cloud representation, then predicts a short motion chunk with a pretrained visual encoder and learned dynamics backbone. Large-scale real/simulated training improves prediction and supports physical manipulation through a separate planner. The central tradeoff is explicit geometric action conditioning versus assumptions about static initial scenes, reliable calibration and accurately realized robot motion.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Motus: A Unified Latent Action World Model

Hongzhe Bi; Hengkai Tan; Shenghao Xie; Zeyuan Wang; Shuhe Huang; Haitian Liu; Ruowen Zhao; Yao Feng; Chendong Xiang; Yinze Rong; Hongyan Zhao; Hanyu Liu; Zhizhong Su; Lei Ma; Hang Su; Jun Zhu

Motus connects video, action and understanding experts through shared attention, then uses separate modality noise levels to switch among five prediction modes. Optical-flow latent actions provide pretraining targets for videos without robot action labels. Execution results favor the complete recipe on average, but do not establish uniform generalization or isolate every architectural contribution.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Vtla: Vision-tactile-language-action model with preference learning for insertion manipulation

Chaofan Zhang; Peng Hao; Xiaoge Cao; Xiaoshuai Hao; Shaowei Cui; Shuo Wang

VTLA converts wrist vision, fingertip contact histories and a peg-specific instruction into corrective insertion actions through Qwen2-VL. Simulation instruction tuning is followed by preferences ranking sampled actions against a reference action. Evidence supports improved OOD action prediction and physical insertion in a matched setup, with small trial counts and incomplete implementation details (e04–e12).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Dyna-2: A 1-million-hour scaling law for world-action models

Dyna Robotics

Dyna-2 tests whether scaling human manipulation video improves robot learning. Its scaling variant co-trains future-video and action prediction through shared representations but remains reactive during control. The source reports improvements in offline robot prediction and in mean normalized robot performance after task post-training. Separate production experiments study VLA comparisons, deployment, language following and fast video generation. Six original HTML visuals explain these findings while preserving their uncertainty and variant boundaries. [e-architecture, e-objective, e-offline, e-onrobot, e-variants]

Resource reviewedIllustrated editionFull text availableRead the report Primary source

Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving

Jianhua Han; Meng Tian; Jiangtong Zhu; Fan He; Huixin Zhang; Sitong Guo; Dechang Zhu; Hao Tang; Pei Xu; Yuze Guo; Minzhe Niu; Haojie Zhu; Qichao Dong; Xuechao Yan; Siyuan Dong; Lu Hou; Qingqiu Huang; Xiaosong Jia; Hang Xu

Percept-WAM strengthens an InternVL2-8B driving model with supervised image-plane and bird's-eye-view perception tokens, then predicts trajectories through action queries. Its strongest planning row reaches 90.2 NAVSIM PDMS; its central evidence concerns perception and trajectory benchmarks, rather than learned future-world rollouts or physical driving execution.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Ctrl-World: A Controllable Generative World Model for Robot Manipulation

Yanjiang Guo; Lucy Xiaoyang Shi; Jianyu Chen; Chelsea Finn

Ctrl-World turns a pretrained video diffusion model into an action-conditioned simulator for existing robot policies. Joint camera prediction, pose-associated history and frame-specific action conditioning support imagined interactions. Human-selected synthetic successes then improve instruction following, while imperfect contact physics limits the simulator's ability to predict actual task completion.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Towards Generalist Embodied AI: A Survey on World Models for VLA Agents

Wentao Tan; Lei Zhu; Bowen Wang; Enci Xie; Baixu Ji; Zengrong Lin; Wenjie Yang; Jingjing Li; Heng Tao Shen

This survey organizes world models for vision-language-action agents by what prediction does: guide a policy, jointly generate actions and futures, supply imitation data, or provide an imagined environment for evaluation and improvement. Its benchmark compilation motivates stronger evaluation, but does not isolate the causal benefit of world modeling.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

MonoDream: Monocular Vision-Language Navigation with Panoramic Dreaming

Shuo Wang; Yongcai Wang; Zhaoxin Fan; Yucheng Wang; Maiyue Chen; Kaihui Wang; Zhizhong Su; Wanting Li; Xudong Cai; Yeying Jin; Deying Li

MonoDream trains a monocular navigation policy to match current and future panoramic RGB/depth features while learning actions and reconstructing instructions. A shared VLM hidden representation carries these objectives; deployment needs only instruction and monocular observations. The evidence concerns simulated navigation improved by auxiliary supervision, not demonstrated panoramic reconstruction or world-model planning at inference.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model

Xingcheng Zhou; Xuyuan Han; Feng Yang; Yunpu Ma; Volker Tresp; Alois Knoll

OpenDriveVLA turns multi-view driving images, textual ego information and commands into six future ego waypoints through a Qwen-based autoregressive VLA. Its distinctive training sequence aligns scene, agent and map tokens to language, teaches driving QA, then forecasts other agents before tuning ego planning. The strongest mechanism evidence is a collision reduction from auxiliary forecasting without a rounded L2 improvement; the evaluation establishes open-loop planning performance, not executed driving safety.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

DriveLaW: Unifying Planning and Video Generation in a Latent Driving World

Tianze Xia; Yongkang Li; Lijun Zhou; Jingfeng Yao; Kaixin Xiong; Haiyang Sun; Bing Wang; Kun Ma; Guang Chen; Hangjun Ye; Wenyu Liu; Xinggang Wang

DriveLaW uses a video diffusion model as the perception backbone for a separate trajectory diffusion model. Features cached during the first video denoising step condition every action denoising step, connecting video pretraining to planning. The reported gains cover nuScenes video generation and NAVSIM planning, but do not prove physical driving safety or guaranteed agreement between generated scenes and trajectories.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation

Tianxing Chen; Zanxin Chen; Baijun Chen; Zijian Cai; Yibin Liu; Zixuan Li; Qiwei Liang; Xianliang Lin; Yiheng Ge; Zhenyu Gu; Weiliang Deng; Yubin Guo; Tian Nian; Xuanbing Xie; Qiangyu Chen; Kailun Su; Tianling Xu; Guodong Liu; Mengkang Hu; Huan-ang Gao; Kaixuan Wang; Zhixuan Liang; Yusen Qin; Xiaokang Yang; Ping Luo; Yao Mu

RoboTwin 2.0 turns task descriptions, annotated objects and manipulation APIs into simulation-tested expert programs, then generates randomized demonstrations for bimanual policy learning. Its value lies in data construction and evaluation: real-world transfer improves in the tested settings, while broad code-generation coverage and reliable visual error diagnosis remain incomplete.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Hierarchical Latent Action Model

Hanjung Kim; Lerrel Pinto; Seon Joo Kim

HiLAM turns observation-only videos into variable-duration latent skills, then uses those skills to pretrain a hierarchical robot policy. Its central evidence is improved simulated LIBERO control, especially with limited demonstrations; qualitative segmentation and frame predictions provide narrower support for the learned representation (e-boundaries, e-policy, e-efficiency, e-qualitative, e-frame-diagnostic).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

DreamDojo: A Real-Time Robot World Model from Large-Scale Human Videos

Shenyuan Gao; William Liang; Kaiyuan Zheng; Ayaan Malik; Seonghyeon Ye; Sihyun Yu; Wei-Cheng Tseng; Yuzhu Dong; Kaichun Mo; Chen-Hsuan Lin; Qianli Ma; Seungjun Nah; Loic Magne; Jiannan Xiang; Yuqi Xie; Ruijie Zheng; Dantong Niu; You Liang Tan; K. R. Zentner; George Kurian; Suneel Indupuru; Pooya Jannaty; Jinwei Gu; Jun Zhang; Jitendra Malik; Pieter Abbeel; Ming-Yu Liu; Yuke Zhu; Joel Jang; Linxi "Jim" Fan

DreamDojo transfers interaction knowledge from human videos into an action-conditioned robot video simulator. Continuous latent actions provide pretraining labels; robot finetuning supplies executable action semantics. A causal distilled student streams faster, with reduced long-horizon image fidelity. Separate policy and value models turn its predicted futures into planning decisions.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation

Zijian Song; Qichang Li; Sihan Qin; Yuhao Chen; Tianshui Chen; Liang Lin; Guangrun Wang

PhysGen adapts NOVA video pretraining into a continuous autoregressive manipulation policy. A shared Transformer predicts visual and action representations, with separate diffusion decoders and an environment feedback loop. Its strongest reported comparison is 90.8% average LIBERO success; real robot performance averages 75%, matching Pi0. These results support useful transfer, while the stronger claim of learning general physical understanding remains unisolated [e-architecture, e-feedback, e-libero, e-real, e-qualitative].

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

WorldGym: World Model as An Environment for Policy Evaluation

Julian Quevedo; Ansh Kumar Sharma; Yixiang Sun; Varad Suryavanshi; Percy Liang; Sherry Yang

WorldGym replaces a hand-built evaluation environment with an action-conditioned video model. A separate robot policy repeatedly acts on generated observations; GPT-4o grades the resulting trajectory. Bridge experiments support useful aggregate policy ranking, while task-level discrepancies, imperfect object physics and reward-model transfer limit how confidently simulated scores predict deployment.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

AutoMoT: A Unified Vision-Language-Action Model with Asynchronous Mixture-of-Transformers for End-to-End Autonomous Driving

Wenhui Huang; Songyan Zhang; Qihang Huang; Zhidong Wang; Zhiqi Mao; Collister Chua; Zhan Chen; Long Chen; Chen Lv

AutoMoT transfers a frozen VLM’s scene representations into a separately trained action expert through layer-wise shared attention. Reusing the understanding expert’s key–value cache lets trajectory prediction run faster than semantic reasoning; an optional diffusion refiner improves simulated driving results. The strongest evidence combines CARLA closed-loop scores with a controlled cache-delay comparison, while real-world data are evaluated open-loop. [e03, e04, e09, e11, e12, e16, e17]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

VLAW: Iterative Co-Improvement of Vision-Language-Action Policy and World Model

Yanjiang Guo; Tony Lee; Lucy Xiaoyang Shi; Jianyu Chen; Percy Liang; Chelsea Finn

VLAW improves a robot policy by first teaching its video world model what the policy actually does, including failures, then learning from reward-filtered imagined successes. Separate π0.5, Ctrl-World, and reward models form an iterative training pipeline. Real-robot results support useful synthetic supervision on five task categories, with unresolved reporting and evaluation details (e-loop, e-policy-results, e-identity).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

ViPRA: Video Prediction for Robot Actions

Sandeep Routray; Hengkai Pan; Unnat Jain; Shikhar Bahl; Deepak Pathak

ViPRA learns discrete motion representations from actionless human and robot videos, jointly pretrains an LWM video-language backbone on future images and latent actions, then adapts that backbone to continuous action chunks. Its control policy does not generate future video at deployment. Gains on SIMPLER and physical manipulation coexist with weaker LIBERO-10 results and unresolved flow-matching notation (e-method, e-pretrain, e-adapt, e-simpler, e-real, e-libero, e-flow-convention).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

Xiaotao Hu; Mingkai Jia; Xiaoyang Guo; Qian Zhang; Xiao-xiao Long; Wei Yin

DrivingWorld predicts driving poses and front-view video through a multimodal autoregressive model. It separates temporal attention from within-frame fusion, then generates pose and image tokens in sequence. Balanced attention protects sparse pose information; corrupted training histories target rollout drift. Reported generation and planning results are promising, but private training data, incomplete implementation details and inconsistent long-video timing limit reproducibility.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

DriveWorld-VLA: Unified Latent-Space World Modeling with Vision-Language-Action for Autonomous Driving

Feiyang jia; Lin Liu; Ziying Song; Caiyan Jia; Hangjun Ye; Xiaoshuai Hao; Long Chen

DriveWorld-VLA connects a driving planner to two BEV prediction branches through shared vision-language hidden states. It first learns scene and trajectory prediction, then learns action-conditioned latent futures, and finally refines the action head using learned reward weights. The reported gains concern benchmark planning scores and collision rates; they do not establish physical deployment or general causal understanding.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation

TRI LBM Team; Jose Barreiros; Andrew Beaulieu; Aditya Bhat; Rick Cory; Eric Cousineau; Hongkai Dai; Ching-Hsin Fang; Kunimatsu Hashimoto; Muhammad Zubair Irshad; Masha Itkina; Naveen Kuppuswamy; Kuan-Hui Lee; Katherine Liu; Dale McConachie; Ian McMahon; Haruki Nishimura; Calder Phillips-Grafflin; Charles Richter; Paarth Shah; Krishnan Srinivasan; Blake Wulfe; Chen Xu; Mengchao Zhang; Alex Alspach; Maya Angeles; Kushal Arora; Vitor Campagnolo Guizilini; Alejandro Castro; Dian Chen; Ting-Sheng Chu; Sam Creasey; Sean Curtis; Richard Denitto; Emma Dixon; Eric Dusel; Matthew Ferreira; Aimee Goncalves; Grant Gould; Damrong Guoy; Swati Gupta; Xuchen Han; Kyle Hatch; Brendan Hathaway; Allison Henry; Hillel Hochsztein; Phoebe Horgan; Shun Iwase; Donovon Jackson; Siddharth Karamcheti; Sedrick Keh; Joseph Masterjohn; Jean Mercat; Patrick Miller; Paul Mitiguy; Tony Nguyen; Jeremy Nimmer; Yuki Noguchi; Reko Ong; Aykut Onol; Owen Pfannenstiehl; Richard Poyner; Leticia Priebe Mendes Rocha; Gordon Richardson; Christopher Rodriguez; Derick Seale; Michael Sherman; Mariah Smith-Jones; David Tago; Pavel Tokmakov; Matthew Tran; Basile Van Hoorick; Igor Vasiljevic; Sergey Zakharov; Mark Zolotas; Rares Ambrus; Kerri Fetzer-Borelli; Benjamin Burchfiel; Hadas Kress-Gazit; Siyuan Feng; Stacie Ford; Russ Tedrake

This study tests whether diverse robot-action pretraining improves specialist manipulation policies. A fixed diffusion-transformer architecture is pretrained on Ramen and finetuned on individual tasks, then compared with single-task training through controlled simulation and blind hardware trials. Benefits are clearest after finetuning and in partial task completion; language steering, preprocessing errors and evaluation protocols limit stronger conclusions.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving

Yingyan Li; Shuyao Shang; Weisong Liu; Bing Zhan; Haochen Wang; Yuqi Wang; Yuntao Chen; Xiaoman Wang; Yasong An; Chufeng Tang; LU HOU; Lue Fan; Zhaoxiang Zhang

DriveVLA-W0 trains driving VLAs to predict visual futures as well as actions, then uses an optional small action expert for efficient trajectory generation. Its strongest controlled evidence is improved large-data planning, while its visual simulator remains a diagnostic capability. Driving inference skips image generation (e-wm, e-scaling, e-expert).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video

Ryan Hoque; Peide Huang; David J. Yoon; Mouli sivapurapu; Jian Zhang

EgoDex pairs egocentric video with tracked human hand poses to support dexterous imitation learning. Its 829 hours and 194 tasks underpin offline hand-trajectory benchmarks: behavior cloning leads for one prediction, while flow matching leads under best-of-many scoring. Robot transfer and world modeling remain proposed uses. [e-dataset, e-comparison, e-usecases]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

DrivingGen: A Comprehensive Benchmark for Generative Video World Models in Autonomous Driving

Yang Zhou; Hao Shao; Letian Wang; Zhuofan Zong; Hongsheng Li; Steven L. Waslander

DrivingGen evaluates driving video generators through complementary visual and recovered-motion measurements. Its two tracks separate broad scenario coverage from ego-trajectory alignment. The central finding is a metric tradeoff: attractive videos, stable agents and accurate commanded motion do not share one universal winner. Trajectory conclusions remain mediated by the reconstruction pipeline.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

WMPO: World Model-based Policy Optimization for Vision-Language-Action Models

Fangqi Zhu; Zhengyang Yan; Zicong Hong; Quanxin Shou; Xiao Ma; Song Guo

WMPO post-trains a vision-language-action policy using complete trajectories generated by a separate action-conditioned video model and judged by a learned success classifier. Its interface to the policy is decoded imagery, although diffusion itself uses VAE latents. It improves manipulation success under matched environment-rollout budgets; model fidelity, reward calibration and substantial training compute remain central qualifications.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

DynVLA: Learning World Dynamics for Action Reasoning in Autonomous Driving

Shuyao Shang; Bing Zhan; Yunfei Yan; Yuqi Wang; Yingyan Li; Yasong An; Xiaoman Wang; Jierui Liu; Lu Hou; Lue Fan; Zhaoxiang Zhang; Tieniu Tan

DynVLA represents driving foresight with discrete ego-motion and environment-motion tokens, then conditions trajectory generation on them. A separately trained tokenizer supplies targets; one autoregressive policy learns dynamics-first action generation and undergoes reinforcement fine-tuning. Reported NAVSIM PDMS reaches 91.7. The tradeoff is compact prediction versus propagation of prediction errors into planning, with benchmark gains but incomplete reproduction details and no demonstrated physical-road deployment (e03–e16, e20–e22).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

DINOv3

Oriane Siméoni; Huy V. Vo; Maximilian Seitzer; Federico Baldassarre; Maxime Oquab; Cijo Jose; Vasil Khalidov; Marc Szafraniec; Seungeun Yi; Michaël Ramamonjisoa; Francisco Massa; Daniel Haziza; Luca Wehrstedt; Jianyuan Wang; Timothée Darcet; Théo Moutakanni; Leonel Sentana; Claire Roberts; Andrea Vedaldi; Jamie Tolan; John Brandt; Camille Couprie; Julien Mairal; Hervé Jégou; Patrick Labatut; Piotr Bojanowski

DINOv3 scales self-supervised image encoding while repairing the dense features that deteriorate during long training. Its Gram anchoring loss transfers patch-similarity structure from an earlier teacher, followed by resolution adaptation and distillation. The resulting frozen features support strong dense probes and trained downstream systems. The contribution is reusable visual representation learning; action selection and predictive environment dynamics are outside the demonstrated system.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Spatial Forcing: Implicit Spatial Representation Alignment for Vision-Language-Action Model

Fuhao Li; Wenxuan Song; Han Zhao; Jingbo Wang; Pengxiang Ding; Donglin Wang; Long Zeng; Haoang Li

Spatial Forcing (SF) fine-tunes a vision-language-action policy with supervision from a pretrained geometry model. It aligns intermediate visual tokens with VGGT features, improving reported manipulation success while retaining the base inference path. The strongest LIBERO average is 98.5%, versus 97.1% for OpenVLA-OFT; component experiments and physical-robot tests support a useful training intervention, with unresolved efficiency and evaluation details (e-alignment, e-libero, e-components, e-real).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Learning Latent Action World Models In The Wild

Quentin Garrido; Tushar Nagarajan; Basile Terver; Nicolas Ballas; Yann LeCun; Michael Rabbat

A frozen video encoder supports jointly learned inverse dynamics and latent-conditioned future prediction. On natural videos, constrained continuous actions capture richer changes than the tested quantization scheme. A separately supervised controller makes this space usable for short-horizon planning, but latent predictability, visual quality and planning accuracy are different objectives.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Uncertainty-Aware Robotic World Model Makes Offline Model-Based Reinforcement Learning Work on Real Robots

Chenhao Li; Andreas Krause; Marco Hutter

RWM-U learns action-conditioned robotic dynamics from fixed logs and uses disagreement among prediction heads to discourage a separately trained PPO policy from exploiting unreliable imagined transitions. Its strongest evidence combines locomotion comparisons, a penalty sweep, and physical ANYmal D data-mixture results. The practical tradeoff is between exploiting useful model predictions and becoming too conservative to move; realism alone also fails when real data replace diverse simulation experience. The reviewed source is the January 2026 arXiv v3 preprint.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Towards Practical World Model-based Reinforcement Learning for Vision-Language-Action Models

Zhilong Zhang; Haoxiang Ren; Yihao Sun; Yifei Sheng; Haonan Wang; Haoxin Lin; Zhichao Wu; Pierre-Luc Bacon; Yang Yu

VLA-MBPO finetunes a separate vision-language-action policy using short imagined rollouts from a Bagel-based dynamics/reward model. Its central choices are action-chunk prediction, head-to-wrist conditional generation and branching from recorded observations. Table 2 reports LIBERO average success rising from 76.8% to 85.9%, with arithmetic discrepancies detailed below; physical-task gains accompany substantial compute costs and unresolved implementation details.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Interactive World Simulator for Robot Policy Training and Evaluation

Yixuan Wang; Rhythm Syed; Fangyu Wu; Mengchao Zhang; Aykut Onol; Jose Barreiros; Hooshang Nayyeri; Tony Dear; Huan Zhang; Yunzhu Li

Interactive World Simulator learns action-conditioned visual dynamics from robot play data, then acts as an interactive environment for collecting demonstrations and evaluating separate policies. Consistency models decode images and predict latent futures. The evidence supports useful simulation within the studied task distributions, with less support for broad equivalence to real demonstrations or unrestricted physical fidelity.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

World Action Models are Zero-shot Policies

Seonghyeon Ye; Yunhao Ge; Kaiyuan Zheng; Shenyuan Gao; Sihyun Yu; George Kurian; Suneel Indupuru; You Liang Tan; Chuning Zhu; Jiannan Xiang; Ayaan Malik; Kyungmin Lee; William Liang; Nadun Ranawaka; Jiasheng Gu; Yinzhen Xu; Guanzhi Wang; Fengyuan Hu; Avnish Narayan; Johan Bjorck; Jing Wang; Gwanghyun Kim; Dantong Niu; Ruijie Zheng; Yuqi Xie; Jimmy Wu; Qi Wang; Ryan Julian; Danfei Xu; Yilun Du; Yevgen Chebotar; Scott Reed; Jan Kautz; Yuke Zhu; Linxi "Jim" Fan; Joel Jang

DreamZero turns a pretrained video diffusion transformer into a policy that jointly denoises visual futures and motor commands, then refreshes its history with real observations. Real-robot results support stronger generalization than the tested VLAs, but single-step acceleration costs some task progress and uses substantial hardware. “Zero-shot” refers to held-out robot tasks or environments, not learning without robot data.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Ground Slow, Move Fast: A Dual-System Foundation Model for Generalizable Vision-and-Language Navigation

Meng Wei; Chenyang Wan; Jiaqi Peng; Xiqian Yu; Yuqiang Yang; Delin Feng; Wenzhe Cai; Chenming Zhu; Tai Wang; Jiangmiao Pang; Xihui Liu

DualVLN separates instruction-grounded pixel-goal planning from fast trajectory generation. A VLM supplies waypoint text and learned latent goals to an RGB-conditioned diffusion policy. Unseen navigation improves, but dynamic-human collisions and incomplete implementation details limit stronger conclusions.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Dream2Flow: Bridging Video Generation and Open-World Manipulation with 3D Object Flow

Karthik Dharmarajan; Wenlong Huang; Jiajun Wu; Li Fei-Fei; Ruohan Zhang

Dream2Flow converts generated human-interaction videos into 3D object trajectories, then uses domain-specific optimization or reinforcement learning to make a robot realize them. Its central benefit is an object-level interface across embodiments; its reliability still depends on video geometry, tracking, contact assumptions and the downstream controller.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Act2Goal: From World Model To General Goal-conditioned Policy

Pengfei Zhou; Liliang Chen; Shengcong Chen; Di Chen; Wenzhi Zhao; Rongjun Jin; Guanghui Ren; Jianlan Luo

Act2Goal converts current and target images into world-model features that guide a separate action expert. Dense near-term control and sparse distal predictions support long-horizon manipulation, with optional hindsight-based LoRA adaptation. Real-robot results and a writing ablation are strong within the tested setups, while sampling, timing and implementation ambiguities constrain reproduction.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

CoVAR: Co-generation of Video and Action for Robotic Manipulation via Multi-Modal Diffusion

Liudi Yang; Yang Bai; George Eskandar; Fengyi Shen; Mohammad Altillawi; Dong Chen; Ziyuan Liu; Abhinav Valada

CoVAR co-generates instruction-conditioned video and robot actions through separate, interacting diffusion transformers. Bridge Attention connects a pretrained video branch to an action branch; a UNet decodes actions, with an additional refinement model for LIBERO90. Reported simulation and UR5 successes support the complete system, while refinement dependence, incomplete evaluation details, and roughly four-second sequence generation limit broader conclusions.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs

Jonas Pai; Liam Achenbach; Victoriano Montesinos; Benedek Forrai; Oier Mees; Elvis Nava

mimic-video turns a robot-adapted video generator into features for a separate action policy. Its default setting combines observed frames with entirely noisy future latents, avoiding video reconstruction while retaining full action denoising. Experiments support strong decoder-data efficiency and manipulation performance, with task-dependent noise tuning and limited deployment evidence (e-architecture, e-sampling, e-efficiency, e-limits).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

World Models for Learning Dexterous Hand-Object Interactions from Human Videos

Raktim Gautam Goswami; Amir Bar; David Fan; Tsung-Yen Yang; Gaoyue Zhou; Prashanth Krishnamurthy; Michael Rabbat; Farshad Khorrami; Yann LeCun

DexWM learns an action-conditioned transition model in frozen visual features, using human hand motion to support dexterous robot planning. A hand-consistency objective preserves local finger information; CEM searches robot joint trajectories through the learned model. Simulation and small-scale physical grasping results support transfer after exploratory simulation fine-tuning, with substantial remaining placement failures and a strict distinction between predicted rollouts and executed actions.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

WorldLens: Full-Spectrum Evaluations of Driving World Models in Real World

Ao Liang; Lingdong Kong; Tianyi Yan; Hongsi Liu; Wesley Yang; Ziqi Huang; Wei Yin; Jialong Zuo; Yixuan Hu; Dekai Zhu; Dongyue Lu; Youquan Liu; Guangfeng Jiang; Linfeng Li; Xiangtai Li; Long Zhuo; Lai Xing Ng; Benoit R. Cottereau; Changxin Gao; Liang Pan; Wei Tsang Ooi; Ziwei Liu

WorldLens evaluates driving video generators through appearance, reconstructability, planner behavior, perception and human judgment. Its strongest lesson is that favorable image metrics coexist with poor closed-loop route completion. A separate LoRA-trained critic learns score-and-rationale outputs from human annotations; its generalization evidence remains qualitative.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

VideoVLA: Video Generators Can Be Generalizable Robot Manipulators

Yichao Shen; Fangyun Wei; Zhiying Du; Yaobo Liang; Yan Lu; Jiaolong Yang; Nanning Zheng; Baining Guo

VideoVLA adapts CogVideoX-5B into a robot policy that jointly denoises future video latents and executable action chunks, conditioned on language and the current image. Its clearest gains concern novel objects and skills transferred between embodiments. Video supervision matters strongly in the reported ablations, but imagined success exceeds executed success and deployment remains slow.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

MIND-V: Hierarchical World Model for Long-Horizon Robotic Manipulation with RL-based Physical Alignment

Ruicheng Zhang; Mingyang Zhang; Jun Zhou; Xiaofan Liu; Zunnan Xu; Zhizhou Zhong; Puxin Yan; Haocheng Luo; Xiu Li

MIND-V generates manipulation videos by converting language into subtasks, grounding object and arm trajectories, and rendering candidate futures with a conditioned diffusion model. RL adjusts a learned physical-coherence proxy; inference searches over generated alternatives. Its strongest evidence concerns video subtask success, with a separate, incompletely specified policy-learning experiment in simulation.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

MindDrive: An All-in-One Framework Bridging World Models and Vision-Language Model for End-to-End Autonomous Driving

Bin Sun; Yaoguang Cao; Yan Wang; Rui Wang; Jiachen Shang; Xiejie Feng; Jiayi Lu; Jia Shi; Shichun Yang; Xiaoyu Yan; Ziying Song

MindDrive couples action-conditioned BEV scene prediction with anchor refinement and a VLM trajectory scorer. Its main contribution is the connection between future-aware candidate generation and selection. NAVSIM scores support planning gains, while incomplete training specifications and inconsistent result reporting limit reproducibility.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Video2Act: A Dual-System Video Diffusion Policy with Robotic Spatio-Motional Modeling

Yueru Jia; Jiaming Liu; Shengbang Liu; Rui Zhou; Wanhe Yu; Yuyang Yan; Xiaowei Chi; Yandong Guo; Boxin Shi; Shanghang Zhang

Video2Act turns observed video into conditioning for a separate robot action policy. Sobel filtering emphasizes spatial structure in video-model features, while temporal Fourier filtering emphasizes motion. A slow video network refreshes these representations for a faster diffusion action head. The strongest evidence is improved executed manipulation on RoboTwin and a small real-robot evaluation; the speed claims require separating action-chunk throughput from fresh-feedback control.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

LatBot: Distilling Universal Latent Actions for Vision-Language-Action Models

Zuolei Li; Xingyu Gao; Xiaofan Wang; Jianlong Fu

LatBot learns scene and motion latents from language-conditioned robot and human manipulation videos, supervises them through future-image and physical-action decoding, then distills them into a VLA student. An action expert converts the student's features into executable commands. The strongest evidence concerns downstream manipulation success; universal physical understanding and preserved reasoning remain broader interpretations of these results (e03–e06, e12–e16).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

GigaWorld-0: World Models as Data Engine to Empower Embodied AI

GigaWorld Team; Angen Ye; Boyuan Wang; Chaojun Ni; Guan Huang; Guosheng Zhao; Haoyun Li; Jiagang Zhu; Kerui Li; Mengyuan Xu; Qiuping Deng; Siting Wang; Wenkang Qin; Xinze Chen; Xiaofeng Wang; Yankai Wang; Yu Cao; Yifan Chang; Yuan Xu; Yun Ye; Yang Wang; Yukun Zhou; Zhengyuan Zhang; Zhehao Dong; Zheng Zhu

GigaWorld-0 produces VLA training data through controllable video generation and a modular 3D simulation pipeline. Dreamer generates observations; a separate inverse-dynamics model supplies action labels. Its strongest measured evidence is a PBench overall-score lead, while DreamGen physical adherence is mixed and downstream robot gains are illustrated without local success-rate tables (e03–e04, e13, e17–e20).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Learning Massively Multitask World Models for Continuous Control

Nicklas Hansen; Hao Su; Xiaolong Wang

Newt extends TD-MPC2 into a language-conditioned agent that learns latent dynamics from demonstrations and online interaction across MMBench. Its practical recipe pretrains the world model and policy prior, retains action supervision, and uses the model to plan. It improves over the evaluated multitask baselines, but specialist policies remain stronger and long open-loop execution is uneven. The evidence concerns simulated, predominantly state-based control.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

RynnVLA-002: A Unified Vision-Language-Action and World Model

Jun Cen; Siteng Huang; Yuqian Yuan; Kehan Li; Hangjie Yuan; Chaohui Yu; Bohan Hou; Yuming Jiang; Jiayan Guo; Xin Li; Hao Luo; Fan Wang; Deli Zhao; Hao Chen

RynnVLA-002 finetunes a shared Chameleon backbone for action prediction and action-conditioned image prediction, then adds a parallel continuous-action head. Its strongest evidence is mutual training benefit: better executed policies and better held-out visual predictions. Policy inference uses no imagined-image rollout. The reported 97.4% LIBERO average is competitive, while physical evidence is limited to SO100 pick-and-place.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Dexterity from Smart Lenses: Multi-Fingered Robot Manipulation with In-the-Wild Human Demonstrations

Irmak Guzey; Haozhi Qi; Julen Urain; Changhao Wang; Jessica Yin; Krishna Bodduluri; Mike Lambeta; Lerrel Pinto; Akshara Rai; Jitendra Malik; Tingfan Wu; Akash Sharma; Homanga Bharadhwaj

AINA converts smart-glass human demonstrations into aligned 3D object and fingertip tracks, learns a future-fingertip policy, and executes it through arm–hand inverse kinematics. Fifty in-the-wild demonstrations and one human demonstration in the deployment scene support task-specific training without robot interaction data. Real-robot transfer is strongest on several simple tasks and substantially weaker on pouring, stowing and knob rotation.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

π^*_0.6: a VLA That Learns From Experience

Physical Intelligence; Ali Amin; Raichelle Aniceto; Ashwin Balakrishna; Kevin Black; Ken Conley; Grace Connors; James Darpinian; Karan Dhabalia; Jared DiCarlo; Danny Driess; Michael Equi; Adnan Esmail; Yunhao Fang; Chelsea Finn; Catherine Glossop; Thomas Godden; Ivan Goryachev; Lachy Groom; Hunter Hancock; Karol Hausman; Gashon Hussein; Brian Ichter; Szymon Jakubczak; Rowan Jen; Tim Jones; Ben Katz; Liyiming Ke; Chandra Kuchi; Marinda Lamb; Devin LeBlanc; Sergey Levine; Adrian Li-Bell; Yao Lu; Vishnu Mano; Mohith Mothukuri; Suraj Nair; Karl Pertsch; Allen Z. Ren; Charvi Sharma; Lucy Xiaoyang Shi; Laura Smith; Jost Tobias Springenberg; Kyle Stachowicz; Will Stoeckle; Alex Swerdlow; James Tanner; Marcel Torne; Quan Vuong; Anna Walling; Haohuan Wang; Blake Williams; Sukwon Yoo; Lili Yu; Ury Zhilinsky; Zhiyuan Zhou

RECAP improves a flow-matching VLA by using a separately trained value function to label recorded actions as relatively advantageous, then conditioning action generation on the positive label. Demonstrations, autonomous trials, and optional human corrections feed repeated offline updates. Real-robot experiments report faster and more reliable laundry folding, espresso preparation, and box assembly, but several recipe and data-accounting inconsistencies limit exact reproduction.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Scalable Policy Evaluation with Video World Models

Wei-Cheng Tseng; Jinwei Gu; Qinsheng Zhang; Hanzi Mao; Ming-Yu Liu; Florian Shkurti; Lin Yen-Chen

This paper turns a pretrained video diffusion model into an action-conditioned simulator for evaluating existing robot policies. A policy acts on generated observations; a vision-language model judges the resulting video to estimate success rates. Policy-rollout augmentation and pretrained weights improve agreement with simulator evaluations. Bridge experiments extend the test to physical performance, but hallucinations, inconsistent camera views and imperfect labels limit reliability. The contribution is a learned evaluation environment, not a jointly learned action-producing controller.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

PAN: A World Model for General, Interactable, and Long-Horizon World Simulation

PAN Team; Jiannan Xiang; Yi Gu; Zihan Liu; Zeyu Feng; Qiyue Gao; Yiyan Hu; Benhao Huang; Guangyi Liu; Yichi Yang; Kun Zhou; Davit Abrahamyan; Arif Ahmad; Ganesh Bannur; Junrong Chen; Kimi Chen; Mingkai Deng; Ruobing Han; Xinqi Huang; Haoqiang Kang; Zheqi Liu; Enze Ma; Hector Ren; Yashowardhan Shinde; Rohan Shingre; Ramsundar Tanikella; Kaiming Tao; Dequan Yang; Xinle Yu; Cong Zeng; Binglin Zhou; Zhengzhong Liu; Zhiting Hu; Eric P. Xing

PAN turns an initial image and successive language actions into simulated video futures. A Qwen-based backbone predicts compact latent states; a Wan-based causal diffusion decoder renders them and returns generated visual history. Its strongest evidence concerns comparative video continuity and prediction; planning uses an external VLM and does not establish physical execution success.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

John Won; Kyungmin Lee; Huiwon Jang; Dongyoung Kim; Jinwoo Shin

DUST augments a frozen vision-language backbone with jointly denoised actions and future visual embeddings. Separate streams exchange information through attention; independent noise levels and unequal sampling budgets accommodate their different dynamics. Controlled comparisons support improved manipulation, but do not establish general causal world understanding.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

π_RL: Online RL Fine-tuning for Flow-based Vision-Language-Action Models

Kang Chen; Zhihao Liu; Tonghe Zhang; Zhen Guo; Si Xu; Hao Lin; Hongzhi Zang; Xiang Li; Quanlu Zhang; Zhaofei Yu; Guoliang Fan; Tiejun Huang; Yu Wang; Chao Yu

πRL makes flow-based robot action policies trainable with online PPO by assigning tractable Gaussian probabilities to denoising transitions. Flow-Noise learns exploration noise; Flow-SDE combines a fixed noise schedule with a nested, optionally shortened MDP. Simulation gains are large, but transfer is stronger for environmental variations than unseen task objectives (e04–e06, e09–e14).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

World Simulation with Video Foundation Models for Physical AI

NVIDIA; :; Arslan Ali; Junjie Bai; Maciej Bala; Yogesh Balaji; Aaron Blakeman; Tiffany Cai; Jiaxin Cao; Tianshi Cao; Elizabeth Cha; Yu-Wei Chao; Prithvijit Chattopadhyay; Mike Chen; Yongxin Chen; Yu Chen; Shuai Cheng; Yin Cui; Jenna Diamond; Yifan Ding; Jiaojiao Fan; Linxi Fan; Liang Feng; Francesco Ferroni; Sanja Fidler; Xiao Fu; Ruiyuan Gao; Yunhao Ge; Jinwei Gu; Aryaman Gupta; Siddharth Gururani; Imad El Hanafi; Ali Hassani; Zekun Hao; Jacob Huffman; Joel Jang; Pooya Jannaty; Jan Kautz; Grace Lam; Xuan Li; Zhaoshuo Li; Maosheng Liao; Chen-Hsuan Lin; Tsung-Yi Lin; Yen-Chen Lin; Huan Ling; Ming-Yu Liu; Xian Liu; Yifan Lu; Alice Luo; Qianli Ma; Hanzi Mao; Kaichun Mo; Seungjun Nah; Yashraj Narang; Abhijeet Panaskar; Lindsey Pavao; Trung Pham; Morteza Ramezanali; Fitsum Reda; Scott Reed; Xuanchi Ren; Haonan Shao; Yue Shen; Stella Shi; Shuran Song; Bartosz Stefaniak; Shangkun Sun; Shitao Tang; Sameena Tasmeen; Lyne Tchapmi; Wei-Cheng Tseng; Jibin Varghese; Andrew Z. Wang; Hao Wang; Haoxiang Wang; Heng Wang; Ting-Chun Wang; Fangyin Wei; Jiashu Xu; Dinghao Yang; Xiaodong Yang; Haotian Ye; Seonghyeon Ye; Xiaohui Zeng; Jing Zhang; Qinsheng Zhang; Kaiwen Zheng; Andrew Zhu; Yuke Zhu

Cosmos-Predict2.5 turns text and optional visual context into video using a flow-matching transformer, then specializes that backbone for spatial control, multiple cameras and supplied robot actions. Cosmos-Transfer2.5 also augments observations for a separately trained robot policy. The strongest execution evidence is a small real-robot robustness experiment; most other evaluations measure generated-video quality, consistency or instruction following. This is a reading of the February 2026 v2 source, not a comparison with its 2025 submission. [e-identity, e-architecture, e-transfer, e-robot-results, e-action]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

GigaBrain-0: A World Model-Powered Vision-Language-Action Model

GigaBrain Team; Angen Ye; Boyuan Wang; Chaojun Ni; Guan Huang; Guosheng Zhao; Haoyun Li; Jie Li; Jiagang Zhu; Lv Feng; Peng Li; Qiuping Deng; Runqi Ouyang; Wenkang Qin; Xinze Chen; Xiaofeng Wang; Yang Wang; Yifan Li; Yilong Li; Yiran Ding; Yuan Xu; Yun Ye; Yukun Zhou; Zhehao Dong; Zhenan Wang; Zhichao Liu; Zheng Zhu

GigaBrain-0 trains a robot VLA on real demonstrations and GigaWorld-generated variants, combining RGBD perception, embodied intermediate predictions and a flow-matching action expert. Controlled data-mixture studies support improved robustness; several prose claims conflict with plotted results. GigaWorld serves as a training-data engine, while Small targets edge execution. [e02, e03, e13, e15, e16, e17, e18, e19]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models

Senyu Fei; Siyin Wang; Junhao Shi; Zihao Dai; Jikun Cai; Pengfang Qian; Li Ji; Xinzhe He; Shiduo Zhang; Zhaoye Fei; Jinlan Fu; Jingjing Gong; Xipeng Qiu

LIBERO-Plus turns high simulation success into a diagnostic question: does a policy still execute the intended task when appearance, geometry, initialization or language changes? Its 10,030-instance benchmark exposes uneven robustness and weak instruction dependence in several settings. Additional OpenVLA-OFT training improves the final benchmark score to 79.5%, while robot-initialization robustness remains poor. Evidence: e02, e05, e10, e15.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Ego-Vision World Model for Humanoid Contact Planning

Hang Liu; Yuman Gao; Sangli Teng; Yufeng Chi; Yakun Sophia Shao; Zhongyu Li; Maani Ghaffari; Koushil Sreenath

An offline-trained recurrent world model predicts the consequences of humanoid posture commands from ego-centric depth and proprioception. Sampling MPC averages learned action-values over short latent rollouts, then sends one command to a separate motor controller. Simulation ablations favor this approach for ball blocking and arch traversal; physical demonstrations establish feasibility with operator-controlled base velocity, rather than autonomous navigation or measured deployment reliability.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Generative World Modelling for Humanoids: 1X World Model Challenge Technical Report

Riccardo Mereu; Aidan Scannell; Yuxin Hou; Yi Zhao; Aditya Jitta; Antonio Dominguez; Luigi Acerbi; Amos Storkey; Paul Chang

Team Revontuli uses two separate, state-conditioned predictors: a LoRA-adapted Wan video model for future RGB frames and a spatio-temporal Transformer for future discrete tokens. Both lead the reported challenge leaderboard, but their scores measure conditional prediction rather than executed humanoid control.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

MimicDreamer: Aligning Human and Robot Demonstrations for Scalable VLA Training

Haoyun Li; Ivan Zhang; Runqi Ouyang; Xiaofeng Wang; Zheng Zhu; Zhiqin Yang; Zhentao Zhang; Boyuan Wang; Chaojun Ni; Wenkang Qin; Xinze Chen; Yun Ye; Guan Huang; Zhenbo Song; Xingang Wang

MimicDreamer converts human demonstrations into robot-looking videos paired with retargeted actions, then post-trains a separate π0 policy. The contribution is training-data alignment across viewpoint, embodiment and appearance. Physical-task results improve with added synthetic demonstrations, but real-robot supervision remains part of the evaluated pipeline and several headline summaries conflict with the tables.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

MoWM: Mixture-of-World-Models for Embodied Planning via Latent-to-Pixel Feature Modulation

Yangcheng Yu; Xin Jin; Yu Shang; Xin Zhang; Haisheng Su; Wei Wu; Yong Li

MoWM conditions an action diffusion decoder on predictions from two frozen world models: an SVD-based video model and a transformer predicting V-JEPA 2 features. Learned concatenation and a pixel-feature residual combine their representations. CALVIN results favor this fusion; decoded prediction images and a real-robot folding example provide narrower qualitative evidence. The report separates those findings from claims about why fusion works.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

LongScape: Advancing Long-Horizon Embodied World Models with Context-Aware MoE

Yu Shang; Lei Jin; Yiding Ma; Xin Zhang; Chen Gao; Wei Wu; Yong Li

LongScape generates embodied manipulation videos by choosing how many frames to predict next. Robot action annotations define variable-length training chunks; four diffusion experts specialize in their lengths, and a text-and-vision router selects an expert during autoregressive rollout. The reported gains concern video similarity and distributional quality on LIBERO and AGIBOT-World. The method supplies neither executed robot actions nor evidence that its videos improve downstream policies.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

World4RL: Diffusion World Models for Policy Refinement with Reinforcement Learning for Robotic Manipulation

Zhennan Jiang; Kai Liu; Yuxin Qin; Shuai Tian; Yupeng Zheng; Mingcai Zhou; Chao Yu; Haoran Li; Dongbin Zhao

World4RL improves an imitation-initialized manipulation policy with PPO inside a frozen diffusion simulator. A separate image-based reward classifier scores predicted states. Two-hot actions, random training rollouts and constrained exploration help keep imagined transitions useful. Simulation gains are clear in the reported table; physical results are promising under fixed initial conditions, but the printed BC aggregate conflicts with its trial counts.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Latent Action Pretraining Through World Modeling

Bahey Tharwat; Yara Nasser; Ali Abouzeid; Ian Reid

LAWM turns next-frame prediction into a pretraining signal for an imitation policy's action chunks. A jointly trained world model gives the policy a video-derived prior, then disappears before supervised robot adaptation. Human-video pretraining improves the reported LIBERO-suite average, but cross-publication comparisons and inconsistent implementation descriptions constrain interpretation.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Spatial Policy: Guiding Visuomotor Robotic Manipulation with Spatial-Aware Modeling and Reasoning

Yijun Liu; Yuwei Liu; Yuan Meng; Jieheng Zhang; Yuwei Zhou; Ye Li; Jiacheng Jiang; Kangye Ji; Shijia Ge; Zhi Wang; Wenwu Zhu

Spatial Policy turns robot–object geometry into a structured plan, conditions video diffusion on that plan, and uses a separate goal-conditioned diffusion policy to execute the imagined trajectory. Video screening and execution feedback close the loop. The strongest evidence is improved simulated task success; the physical demonstration remains limited.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

MoVieDrive: Urban Scene Synthesis with Multi-Modal Multi-View Video Diffusion Transformer

Guile Wu; David Huang; Dongfeng Bai; Bingbing Liu

MoVieDrive jointly synthesizes RGB, depth, and semantic videos across driving cameras using shared diffusion layers and modality-specific interaction. Layouts and optional initial frames condition generation. Its experiments support video synthesis and auxiliary-modality quality, with an RGB-quality tradeoff; they do not demonstrate executed driving control (e-overview, e-modality-ablation, e-limitations).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model

Anqing Jiang; Yu Gao; Yiru Wang; Zhigang Sun; Shuo Wang; Yuwen Heng; Hao Sun; Shichen Tang; Lijuan Zhu; Jinhao Chai; Jijun Wang; Zichong Gu; Hao Jiang; Li Sun

IRL-VLA trains an autonomous-driving diffusion policy with a separate model of trajectory rewards. Imitation initializes the policy, simulator-labeled trajectories train the reward model, and reinforcement learning refines the policy. The reported Navhard gain is small in aggregate and redistributes progress, safety and comfort scores. The evidence concerns simulated planning rather than physical driving.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation

Yue Liao; Pengfei Zhou; Siyuan Huang; Donglin Yang; Shengcong Chen; Yuxin Jiang; Yue Hu; Jingbin Cai; Si Liu; Jianlan Luo; Liliang Chen; Shuicheng Yan; Maoqing Yao; Guanghui Ren

Genie Envisioner transfers robotic video-generation representations into two downstream systems: GE-Act decodes executable action chunks from visual latents, while GE-Sim predicts videos conditioned on supplied actions. Its controlled ablation supports embodied pretraining, with state-dependent benefits from task adaptation. Cross-embodiment results require demonstrations and a new action head; simulator video scores do not establish physical policy success.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Video Generators are Robot Policies

Junbang Liang; Pavel Tokmakov; Ruoshi Liu; Sruthi Sudhakar; Paarth Shah; Rares Ambrus; Carl Vondrick

Video Policy turns Stable Video Diffusion into a manipulation policy through a separate action diffusion U-Net conditioned on intermediate video features. Its main 50-demonstration RoboCasa configuration trains video prediction first and then freezes that representation for action learning. Simulation results support this approach, while sparse physical trials and slow video inference limit broader conclusions.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Vidar: Embodied Video Diffusion Model for Generalist Manipulation

Yao Feng; Hengkai Tan; Xinyi Mao; Chendong Xiang; Guodong Liu; Shuhe Huang; Hang Su; Jun Zhu

Vidar adapts an embodied video generator to a target bimanual robot, then converts predicted imagery into controls using a separately trained masked inverse dynamics model. Low-demonstration real-world results and component ablations support this design; action observability, open-loop execution, and substantial prior training constrain its generality.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling

Meng Wei; Chenyang Wan; Xiqian Yu; Tai Wang; Yuqiang Yang; Xiaohan Mao; Chenming Zhu; Wenzhe Cai; Hanqing Wang; Yilun Chen; Xihui Liu; Jiangmiao Pang

StreamVLN turns a Video-LLM into a streaming navigation policy: it reuses recent dialogue KV states and compresses older visual context into memory. Optional geometric pruning removes repeated spatial observations. Simulation, robot and latency experiments support the design, but inconsistent prose/table scores and missing implementation details limit precise reproduction.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Critique of World Model

Eric Xing; Mingkai Deng; Jinyu Hou

This theoretical essay defines a world model through its role in simulating action-dependent futures for a separate decision-making agent. It critiques five design choices and proposes Generative Latent Prediction (GLP), instantiated as Physical, Agentic, and Nested (PAN): hierarchical discrete concepts and continuous embeddings, an enhanced LLM plus diffusion predictor, and observation-grounded decoding. Its contribution is a design argument and mathematical analysis; this PDF supplies no PAN benchmark results. The stronger loss guarantees require qualifications discussed below.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge

Wenyao Zhang; Hongsi Liu; Zekun Qi; Yunnan Wang; Xinqiang Yu; Jiazhao Zhang; Runpei Dong; Jiawei He; Fan Lu; He Wang; Zhizheng Zhang; Li Yi; Wenjun Zeng; Xin Jin

DreamVLA trains a shared backbone to anticipate motion-relevant regions, depth and semantic features, then conditions action diffusion on its latent predictions. Explicit visual decoders disappear at inference. It reports 4.44 completed instructions on CALVIN ABC-D and 76.7% real-robot success under an attempt-limited protocol. Its lesson is selective forecasting, tempered by unresolved mask, configuration and table inconsistencies.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

RoboScape: Physics-informed Embodied World Model

Yu Shang; Xin Zhang; Yinzhou Tang; Lei Jin; Chen Gao; Wei Wu; Yong Li

RoboScape predicts action-conditioned robotic RGB and depth videos using coupled autoregressive branches. Depth-feature feedback and motion-selected keypoint supervision aim to improve geometry and interaction dynamics. Reported benefits include stronger video metrics, useful synthetic policy-training data, and correlated simulator-based policy evaluation; the ablations expose tradeoffs, and physical robot deployment remains future work.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

WorldVLA: Towards Autoregressive Action World Model

Jun Cen; Chaohui Yu; Hangjie Yuan; Yuming Jiang; Siteng Huang; Jiayan Guo; Xin Li; Yibing Song; Hao Luo; Fan Wang; Deli Zhao; Hao Chen

WorldVLA adapts Chameleon into one token-generating model with two tasks: produce robot actions from images and instructions, or predict images from observations and actions. Joint training improves LIBERO policy success and longer-horizon video prediction in the reported ablations. Its distinctive attention mask removes dependencies between successive actions in a chunk. The evidence supports shared-model learning benefits, but does not establish an inference-time planner that ranks imagined futures.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

AMPLIFY: Actionless Motion Priors for Robot Learning from Videos

Jeremy A. Collins; Loránd Cheng; Kunal Aneja; Albert Wilcox; Benjamin Joffe; Animesh Garg

AMPLIFY learns a compact vocabulary of visual motion from point tracks, predicts that motion from an image and task instruction, and translates it into robot actions through a separate inverse model. Its strongest evidence concerns scarce target-task action labels, including transfer where target-task videos remain available. Better track prediction and physical task completion are evaluated separately.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Mido Assran; Adrien Bardes; David Fan; Quentin Garrido; Russell Howes; Mojtaba; Komeili; Matthew Muckley; Ammar Rizvi; Claire Roberts; Koustuv Sinha; Artem Zholus; Sergio Arnaud; Abha Gejji; Ada Martin; Francois Robert Hogan; Daniel Dugas; Piotr Bojanowski; Vasil Khalidov; Patrick Labatut; Francisco Massa; Marc Szafraniec; Kapil Krishnakumar; Yong Li; Xiaodong Ma; Sarath Chandar; Franziska Meier; Yann LeCun; Michael Rabbat; Nicolas Ballas

V-JEPA 2 learns video features through masked representation prediction. A separate action-conditioned predictor uses the frozen encoder to support image-goal planning on real robot arms. Classification probes, anticipation probes and language alignment test complementary capabilities; robot control remains dependent on subgoals and camera placement.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Seedance 1.0: Exploring the Boundaries of Video Generation Models

Yu Gao; Haoyuan Guo; Tuyen Hoang; Weilin Huang; Lu Jiang; Fangyuan Kong; Huixia Li; Jiashi Li; Liang Li; Xiaojie Li; Xunsong Li; Yifu Li; Shanchuan Lin; Zhijie Lin; Jiawei Liu; Shu Liu; Xiaonan Nie; Zhiwu Qing; Yuxi Ren; Li Sun; Zhi Tian; Rui Wang; Sen Wang; Guoqiang Wei; Guohong Wu; Jie Wu; Ruiqi Xia; Fei Xiao; Xuefeng Xiao; Jiangqiao Yan; Ceyuan Yang; Jianchao Yang; Runkai Yang; Tao Yang; Yihang Yang; Zilyu Ye; Xuejiao Zeng; Yan Zeng; Heng Zhang; Yang Zhao; Xiaozheng Zheng; Peihao Zhu; Jiaxin Zou; Feilong Zuo

Seedance 1.0 combines a bilingual video diffusion backbone with dense captioning, a learned resolution refiner, human preference alignment, and inference acceleration. Its spatial blocks mix language and vision; temporal blocks propagate visual information across frames. A common conditioning interface supports text-to-video and image-to-video, while interleaved shot captions support multi-shot sequences. The reported preference results favor the overall system, but do not isolate its components or establish an action-executing world model.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

3DFlowAction: Learning Cross-Embodiment Manipulation from 3D Flow World Model

Hongyan Zhi; Peihao Chen; Siyuan Zhou; Yubo Dong; Quanxi Wu; Lei Han; Mingkui Tan

3DFlowAction learns instruction-conditioned object trajectories from human and robot videos, checks a rendered endpoint with GPT-4o, and converts accepted flow into robot poses through grasp selection and optimization. Its four-task physical evaluation reports 70% success, but transfer depends on rigid grasp geometry and follows task-specific human-video fine-tuning.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction

Chenyou Fan; Fangzheng Yan; Chenjia Bai; Jiepeng Wang; Chi Zhang; Zhen Wang; Xuelong Li

CogRobot adapts video diffusion to bimanual control through three learned components: instruction-conditioned flow prediction, flow-conditioned RGB prediction, and a task-specific goal-reaching action policy. Flow offers an intermediate description of motion without requiring action labels for the video models. The strongest physical result is Pull Box success of 0.75 versus 0.05 for DP, but the evidence does not establish a universal action policy or broad unseen-task transfer (e02-framing, e05-video, e06-controller, e09-real, e13-limit).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

RoboTransfer: Controllable Geometry-Consistent Video Diffusion for Manipulation Policy Transfer

Liu Liu; Xiaofeng Wang; Guosheng Zhao; Keyu Li; Wenkang Qin; Jiagang Zhu; Jiaxiong Qiu; Zheng Zhu; Guan Huang; Zhizhong Su

RoboTransfer augments robot demonstrations by changing their visual appearance while conditioning video diffusion on the demonstrated geometry. Jointly encoded camera views, metric depth, normals and separate background/object references support multi-view synthesis. A separately trained ACT policy benefits from the augmented observations on two physical manipulation tasks. This is evidence for offline data augmentation, with remaining uncertainty about statistical reliability and physical fidelity.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Learning Generalizable Robot Policy with Human Demonstration Video as a Prompt

Xiang Zhu; Yichen Liu; Hezhong Li; Jianyu Chen

A human demonstration video conditions a frozen video model whose features guide a separate diffusion policy for an Xhand robot. Human hand motions are reconstructed and retargeted into the policy's action space, allowing human demonstrations to supervise action learning. The reported transfer concerns objects and skills absent from robot training demonstrations but present in human training data; it does not establish learning an entirely untrained skill from the inference prompt alone.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

WorldEval: World Model as Real-World Robot Policies Evaluator

Yaxuan Li; Yichen Zhu; Junjie Wen; Chaomin Shen; Yi Xu

WorldEval estimates robot-policy rankings by turning internal policy embeddings into generated manipulation videos and judging their outcomes. Policy2Vec conditions a separately adapted WAN 2.1 simulator; Gemini-2.0 supplies success labels. Paired experiments support relative-ranking usefulness on the tested tabletop setup, while imperfect action fidelity, checkpoint-distribution dependence and inconsistent source reporting limit stronger conclusions.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

ProphetDWM: A Driving World Model for Rolling Out Future Actions and Videos

Xiaodong Wang; Peixi Peng

ProphetDWM predicts driving actions and videos from observations and a short supplied action sequence. An MLP forecasts actions and exposes latent action features to a diffusion video model; a separate visual-context pathway anchors generation. Joint training improves reported nuScenes video metrics, while action prediction improves average L1 error but does not lead on steering alone. The evidence concerns prediction and generated rollouts, with deployment and evaluation-protocol gaps remaining [E03–E10, E17].

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

FLARE: Robot Learning with Implicit World Modeling

Ruijie Zheng; Jing Wang; Scott Reed; Johan Bjorck; Yu Fang; Fengyuan Hu; Joel Jang; Kaushil Kundalia; Zongyu Lin; Loic Magne; Avnish Narayan; You Liang Tan; Guanzhi Wang; Qi Wang; Jiannan Xiang; Yinzhen Xu; Seonghyeon Ye; Jan Kautz; Furong Huang; Yuke Zhu; Linxi Fan

FLARE trains a flow-matching robot policy to match future observation embeddings inside its action-denoising transformer. Compact, action-trained visual-language targets supply an auxiliary learning signal, including from action-free human videos. Simulation and real-robot improvements support this training recipe; they do not establish an explicit planner or calibrated world simulator.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

EnerVerse-AC: Envisioning Embodied Environments with Action Condition

Yuxin Jiang; Shengcong Chen; Siyuan Huang; Liliang Chen; Pengfei Zhou; Yue Liao; Xindong He; Chiming Liu; Hongsheng Li; Maoqing Yao; Guanghui Ren

EVAC turns robot action sequences into future camera observations using a video diffusion model. Spatial pose maps, temporal action differences and camera rays condition the generated environment. A separate policy can interact with that environment for evaluation, or learn from synthetic trajectories. The strongest evidence is a small policy-data augmentation experiment and agreement with real-robot evaluation trends; physical accuracy remains incompletely measured.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models

Yue Hu; Siyuan Huang; Yue Liao; Shengcong Chen; Pengfei Zhou; Liliang Chen; Maoqing Yao; Guanghui Ren

EWMBench evaluates instruction-conditioned robot videos through scene consistency, end-effector motion, and semantics. Its seven-model comparison favors domain-adapted generators, while controlled trajectory corruptions expose why static-looking plausibility is insufficient. These are offline video-evaluation results, with unresolved reporting inconsistencies, rather than demonstrations of executed robot control.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

UniVLA: Learning to Act Anywhere with Task-centric Latent Actions

Qingwen Bu; Yanting Yang; Jisong Cai; Shenyuan Gao; Guanghui Ren; Maoqing Yao; Ping Luo; Hongyang Li

UniVLA learns a discrete action vocabulary from language-annotated videos, teaches a vision-language policy to predict that vocabulary, and adapts a visual-conditioned head to physical controls. Its distinguishing idea is to separate task-relevant motion from distracting changes before policy pretraining. Strong manipulation and navigation results support transfer, but deployment still needs action-labeled adaptation; future-observation prediction trains the latent representation rather than serving as the evaluated inference-time planner (e-lam, e-policy, e-decode, e-libero, e-nav, e-limitations).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

CLAM: Continuous Latent Action Models for Robot Learning from Unlabeled Demonstrations

Anthony Liang; Pavel Czempin; Matthew M. Hong; Yutai Zhou; Jingzhen Wang; Erdem Bıyık; Stephen Tu

CLAM learns continuous action codes from robot observation transitions, grounds them with limited action-labeled data, then imitates expert videos in that latent space. Deployment uses a policy and action decoder; future-observation prediction serves training. Its strongest evidence concerns action-label scarcity within one robot embodiment, with several unresolved protocol details.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Learned Perceptive Forward Dynamics Model for Safe and Platform-aware Robotic Navigation

Pascal Roth; Jonas Frey; Cesar Cadena; Marco Hutter

A perceptive forward dynamics model predicts how a particular robot and locomotion policy will respond to velocity commands on nearby terrain. A separate MPPI planner searches those predictions for useful, low-risk motion. Simulation comparisons support improved navigation, while real-world data improves trajectory prediction; neither establishes universal safety.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

PIN-WM: Learning Physics-INformed World Models for Non-Prehensile Manipulation

Wenxuan Li; Hang Zhao; Zhiyuan Yu; Yu Du; Qin Zou; Ruizhen Hu; Kai Xu

PIN-WM fits an object-specific rigid-body simulator from multiview images and simple interaction video, then trains a separate visual manipulation policy in nearby randomized simulators. Its strongest evidence is real robot transfer for pushing and flipping; the central tradeoff is dependence on known geometry, visual alignment and restrictive physical assumptions.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets

Chuning Zhu; Raymond Yu; Siyuan Feng; Benjamin Burchfiel; Paarth Shah; Abhishek Gupta

Unified World Models (UWM) couples action and future-image denoising in one transformer, with independent noise levels for each modality. Dynamics supervision and action-free video improve finetuned robot policies. The same network supports policy, video, forward-dynamics and inverse-dynamics inference, but the experiments do not establish reliable model-based planning.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models

Qingqing Zhao; Yao Lu; Moo Jin Kim; Zipeng Fu; Zhuoyang Zhang; Yecheng Wu; Zhaoshuo Li; Qianli Ma; Song Han; Chelsea Finn; Ankur Handa; Ming-Yu Liu; Donglai Xiang; Gordon Wetzstein; Tsung-Yi Lin

CoT-VLA makes a 7B multimodal model generate a future subgoal image before predicting a robot action chunk. The shared model learns image prediction from demonstrations and captioned videos, while action supervision requires demonstrations. Experiments support useful goal conditioning and stronger average manipulation performance, with uneven task gains, substantial inference overhead and reporting inconsistencies. Visual chain of thought here means an explicit predicted image used for control; general reasoning is not independently established.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving

Lloyd Russell; Anthony Hu; Lorenzo Bertoni; George Fedoseev; Jamie Shotton; Elahe Arani; Gianluca Corrado

GAIA-2 generates driving video from supplied ego-motion, agent, scene and camera conditions. Independently encoded views enter a shared flow-matching transformer, whose predicted latents decode into surround-view video. The main contribution is a flexible simulator interface for generation, rollouts and editing. Its evidence consists of qualitative examples and improving internal validation curves; it does not establish a downstream driving-safety gain or a learned action-execution policy (overview, inference, curves).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Wan: Open and Advanced Large-Scale Video Generative Models

Team Wan; Ang Wang; Baole Ai; Bin Wen; Chaojie Mao; Chen-Wei Xie; Di Chen; Feiwu Yu; Haiming Zhao; Jianxiao Yang; Jianyuan Zeng; Jiayu Wang; Jingfeng Zhang; Jingren Zhou; Jinkai Wang; Jixuan Chen; Kai Zhu; Kang Zhao; Keyu Yan; Lianghua Huang; Mengyang Feng; Ningyi Zhang; Pandeng Li; Pingyu Wu; Ruihang Chu; Ruili Feng; Shiwei Zhang; Siyang Sun; Tao Fang; Tianxing Wang; Tianyi Gui; Tingyu Weng; Tong Shen; Wei Lin; Wei Wang; Wei Wang; Wenmeng Zhou; Wente Wang; Wenting Shen; Wenyuan Yu; Xianzhong Shi; Xiaoming Huang; Xin Xu; Yan Kou; Yangyu Lv; Yifei Li; Yijing Liu; Yiming Wang; Yingya Zhang; Yitong Huang; Yong Li; You Wu; Yu Liu; Yulin Pan; Yun Zheng; Yuntao Hong; Yupeng Shi; Yutong Feng; Zeyinzi Jiang; Zhen Han; Zhi-Fan Wu; Ziyu Liu

Wan is a video-generation foundation-model family combining a causal video VAE, a text-conditioned diffusion transformer and large-scale curated image/video training. Its strongest evidence is comparative generation quality; image, editing, identity, camera, streaming and audio extensions broaden media creation without establishing an executed-action policy.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Gemma 3 Technical Report

Gemma Team; Aishwarya Kamath; Johan Ferret; Shreya Pathak; Nino Vieillard; Ramona Merhej; Sarah Perrin; Tatiana Matejovicova; Alexandre Ramé; Morgane Rivière; Louis Rouillard; Thomas Mesnard; Geoffrey Cideron; Jean-bastien Grill; Sabela Ramos; Edouard Yvinec; Michelle Casbon; Etienne Pot; Ivo Penchev; Gaël Liu; Francesco Visin; Kathleen Kenealy; Lucas Beyer; Xiaohai Zhai; Anton Tsitsulin; Robert Busa-Fekete; Alex Feng; Noveen Sachdeva; Benjamin Coleman; Yi Gao; Basil Mustafa; Iain Barr; Emilio Parisotto; David Tian; Matan Eyal; Colin Cherry; Jan-Thorsten Peter; Danila Sinopalnikov; Surya Bhupatiraju; Rishabh Agarwal; Mehran Kazemi; Dan Malkin; Ravin Kumar; David Vilar; Idan Brusilovsky; Jiaming Luo; Andreas Steiner; Abe Friesen; Abhanshu Sharma; Abheesht Sharma; Adi Mayrav Gilady; Adrian Goedeckemeyer; Alaa Saade; Alex Feng; Alexander Kolesnikov; Alexei Bendebury; Alvin Abdagic; Amit Vadi; András György; André Susano Pinto; Anil Das; Ankur Bapna; Antoine Miech; Antoine Yang; Antonia Paterson; Ashish Shenoy; Ayan Chakrabarti; Bilal Piot; Bo Wu; Bobak Shahriari; Bryce Petrini; Charlie Chen; Charline Le Lan; Christopher A. Choquette-Choo; CJ Carey; Cormac Brick; Daniel Deutsch; Danielle Eisenbud; Dee Cattle; Derek Cheng; Dimitris Paparas; Divyashree Shivakumar Sreepathihalli; Doug Reid; Dustin Tran; Dustin Zelle; Eric Noland; Erwin Huizenga; Eugene Kharitonov; Frederick Liu; Gagik Amirkhanyan; Glenn Cameron; Hadi Hashemi; Hanna Klimczak-Plucińska; Harman Singh; Harsh Mehta; Harshal Tushar Lehri; Hussein Hazimeh; Ian Ballantyne; Idan Szpektor; Ivan Nardini; Jean Pouget-Abadie; Jetha Chan; Joe Stanton; John Wieting; Jonathan Lai; Jordi Orbay; Joseph Fernandez; Josh Newlan; Ju-yeong Ji; Jyotinder Singh; Kat Black; Kathy Yu; Kevin Hui; Kiran Vodrahalli; Klaus Greff; Linhai Qiu; Marcella Valentine; Marina Coelho; Marvin Ritter; Matt Hoffman; Matthew Watson; Mayank Chaturvedi; Michael Moynihan; Min Ma; Nabila Babar; Natasha Noy; Nathan Byrd; Nick Roy; Nikola Momchev; Nilay Chauhan; Noveen Sachdeva; Oskar Bunyan; Pankil Botarda; Paul Caron; Paul Kishan Rubenstein; Phil Culliton; Philipp Schmid; Pier Giuseppe Sessa; Pingmei Xu; Piotr Stanczyk; Pouya Tafti; Rakesh Shivanna; Renjie Wu; Renke Pan; Reza Rokni; Rob Willoughby; Rohith Vallu; Ryan Mullins; Sammy Jerome; Sara Smoot; Sertan Girgin; Shariq Iqbal; Shashir Reddy; Shruti Sheth; Siim Põder; Sijal Bhatnagar; Sindhu Raghuram Panyam; Sivan Eiger; Susan Zhang; Tianqi Liu; Trevor Yacovone; Tyler Liechty; Uday Kalra; Utku Evci; Vedant Misra; Vincent Roseberry; Vlad Feinberg; Vlad Kolesnikov; Woohyun Han; Woosuk Kwon; Xi Chen; Yinlam Chow; Yuvein Zhu; Zichuan Wei; Zoltan Egyed; Victor Cotruta; Minh Giang; Phoebe Kirk; Anand Rao; Kat Black; Nabila Babar; Jessica Lo; Erica Moreira; Luiz Gustavo Martins; Omar Sanseviero; Lucas Gonzalez; Zach Gleicher; Tris Warkentin; Vahab Mirrokni; Evan Senter; Eli Collins; Joelle Barral; Zoubin Ghahramani; Raia Hadsell; Yossi Matias; D. Sculley; Slav Petrov; Noah Fiedel; Noam Shazeer; Oriol Vinyals; Jeff Dean; Demis Hassabis; Koray Kavukcuoglu; Clement Farabet; Elena Buchatskaya; Jean-Baptiste Alayrac; Rohan Anil; Dmitry; Lepikhin; Sebastian Borgeaud; Olivier Bachem; Armand Joulin; Alek Andreev; Cassidy Hardin; Robert Dadashi; Léonard Hussenot

Gemma 3 combines a distilled language decoder, pooled SigLIP image tokens and mostly local attention to expand visual and long-context capabilities within a smaller inference-memory budget. Strong instruction-tuned math results coexist with substantial long-context accuracy loss and incomplete training disclosure (architecture, vision, it-benchmarks, long-results).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

NVIDIA; :; Johan Bjorck; Fernando Castañeda; Nikita Cherniadev; Xingye Da; Runyu Ding; Linxi "Jim" Fan; Yu Fang; Dieter Fox; Fengyuan Hu; Spencer Huang; Joel Jang; Zhenyu Jiang; Jan Kautz; Kaushil Kundalia; Lawrence Lao; Zhiqi Li; Zongyu Lin; Kevin Lin; Guilin Liu; Edith Llontop; Loic Magne; Ajay Mandlekar; Avnish Narayan; Soroush Nasiriany; Scott Reed; You Liang Tan; Guanzhi Wang; Zu Wang; Jing Wang; Qi Wang; Jiannan Xiang; Yuqi Xie; Yinzhen Xu; Zhenjia Xu; Seonghyeon Ye; Zhiding Yu; Ao Zhang; Hao Zhang; Yizhou Zhao; Ruijie Zheng; Yuke Zhu

GR00T N1 couples Eagle-2 vision-language features to an embodiment-aware flow-matching action transformer. Human videos, simulated demonstrations and generated robot videos expand its training supervision. The strongest evidence is improved post-training performance on tabletop manipulation; the paper does not establish general humanoid locomotion or inference-time world-model planning. [e02, e03, e14, e18]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control

NVIDIA; :; Hassan Abu Alhaija; Jose Alvarez; Maciej Bala; Tiffany Cai; Tianshi Cao; Liz Cha; Joshua Chen; Mike Chen; Francesco Ferroni; Sanja Fidler; Dieter Fox; Yunhao Ge; Jinwei Gu; Ali Hassani; Michael Isaev; Pooya Jannaty; Shiyi Lan; Tobias Lasser; Huan Ling; Ming-Yu Liu; Xian Liu; Yifan Lu; Alice Luo; Qianli Ma; Hanzi Mao; Fabio Ramos; Xuanchi Ren; Tianchang Shen; Xinglong Sun; Shitao Tang; Ting-Chun Wang; Jay Wu; Jiashu Xu; Stella Xu; Kevin Xie; Yuchong Ye; Xiaodong Yang; Xiaohui Zeng; Yu Zeng

Cosmos-Transfer1 turns structured condition videos into visually varied RGB videos using separately trained control branches around a frozen diffusion transformer. Spatial and temporal weights decide which constraints dominate each region. Its experiments establish a controllability–diversity tradeoff and video-generation throughput; practical policy-transfer benefits remain untested.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

LUMOS: Language-Conditioned Imitation Learning with World Models

Iman Nematollahi; Branton DeMoss; Akshay L Chandra; Nick Hawes; Wolfram Burgard; Ingmar Posner

LUMOS trains a language-guided manipulation policy by practicing inside a world model learned from offline robot play. Frozen recurrent dynamics support actor-critic learning with expert-latent matching rewards; hindsight plans and language alignment organize behavior. CALVIN and physical tabletop experiments support this combination, with modest gains over adapted HULC and larger component-ablation losses. Transfer means deployment after learning from that environment’s offline data, without online policy fine-tuning.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Unified Video Action Model

Shuang Li; Yihuai Gao; Dorsa Sadigh; Shuran Song

UVA learns shared video–action latents, then decodes future images and action chunks through separate diffusion heads. Policy deployment skips video decoding while retaining joint-training supervision. Multitask gains coexist with single-task losses and deployment latency tradeoffs.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Moo Jin Kim; Chelsea Finn; Percy Liang

OpenVLA-OFT adapts a pretrained robot policy by predicting continuous action chunks in one decoder pass and training with L1 regression. OFT+ additionally conditions visual features on language through FiLM. The experiments support faster action generation and strong task performance under focused demonstrations, while leaving multimodal behavior and broad pretraining unresolved. The headline 97.1% LIBERO success and 26× throughput describe different input configurations (evidence: libero-success, libero-speed).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

VaViM and VaVAM: Autonomous Driving through Video Generative Modeling

Florent Bartoccioni; Elias Ramzi; Victor Besnier; Shashanka Venkataramanan; Tuan-Hung Vu; Yihong Xu; Loick Chambon; Spyros Gidaris; Serkan Odabas; David Hurych; Renaud Marlet; Alexandre Boulch; Mickael Chen; Éloi Zablocki; Andrei Bursuc; Eduardo Valle; Matthieu Cord

VaViM learns camera-history representations by predicting discrete video tokens; VaVAM adds a flow-matching expert that generates driving trajectories. Larger models improve image-distribution fidelity and best-of-five trajectory matching, yet closed-loop collision avoidance does not improve consistently. This is a study of generative representation transfer to reactive imitation, with simulated driving evaluation and an explicit gap to predictive planning. [e-overview, e-open, e-diagnostic, e-limits]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

The Role of World Models in Shaping Autonomous Driving: A Comprehensive Survey

Sifan Tu; Xin Zhou; Dingkang Liang; Xingyu Jiang; Yumeng Zhang; Xiaofan Li; Xiang Bai

This survey organizes driving world models by what they predict and how those predictions enter autonomous driving. Its useful distinction is between reconstructing future sensory scenes, learning task-oriented latent dynamics, and simulating traffic behavior. These representations can support simulation, synthetic data, planning or pre-training, but a good generation score does not by itself establish reliable closed-loop driving. The report reads the February 2026 revision and treats its benchmark tables as literature compilations.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

VILP: Imitation Learning with Latent Video Planning

Zhengtong Xu, Qiang Qiu, Yu She

VILP learns observation-conditioned future videos in a compressed latent space, decodes them, and maps adjacent frames to actions through a separate low-level policy. This makes short-horizon video planning practical on the tested tasks and lets task videos supply information beyond scarce action labels. Simulation gains depend on data and evaluation protocol; real-robot evidence comprises 15 trials per method.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos

Zhongwei Ren; Yunchao Wei; Xun Guo; Yao Zhao; Bingyi Kang; Jiashi Feng; Xiaojie Jin

VideoWorld predicts compact multi-step dynamics codes alongside video frames, then converts predictions into task operations. On rendered 9×9 Go and simulated robot tasks, this representation substantially improves over video-only prediction. Expert-curated training data, language conditioning and a separately supervised robot action decoder qualify the broader claim of learning solely from unlabeled videos.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

FAST: Efficient Action Tokenization for Vision-Language-Action Models

Karl Pertsch; Kyle Stachowicz; Brian Ichter; Danny Driess; Suraj Nair; Quan Vuong; Oier Mees; Chelsea Finn; Sergey Levine

FAST makes continuous robot action chunks easier to learn with next-token prediction by converting them into quantized frequency coefficients and compressing those coefficients with byte-pair encoding. Its main benefit is efficient policy training on smooth, high-frequency action data. The paper demonstrates executed manipulation skills and reports faster training than a flow-matching VLA, while acknowledging substantially slower autoregressive inference. FAST+ reuses a tokenizer learned across embodiments; this is action representation learning, with no predicted visual world state.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

RoboHorizon: An LLM-Assisted Multi-View World Model for Long-Horizon Robotic Manipulation

Zixuan Chen; Jing Huo; Yangtao Chen; Yang Gao

RoboHorizon trains a visual manipulation policy using LLM-generated staged rewards, representations learned from intervals between demonstration keyframes, and a separate latent dynamics model. The policy learns through imagined rollouts and expert actions, then controls from a front camera. The reported simulation gains are substantial, but conflicting data counts and an apparent reward-code error complicate reproduction.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation

Siyuan Huang; Liliang Chen; Pengfei Zhou; Shengcong Chen; Yue Liao; Zhengkai Jiang; Yue Hu; Peng Gao; Hongsheng Li; Maoqing Yao; Guanghui Ren

EnerVerse learns instruction-conditioned, multi-view future-video representations, then conditions a separate action diffusion head on their backbone features. Sparse history supports long tasks; an offline 4D Gaussian Splatting loop refines training videos. Its strongest benchmark average uses depth-rendered auxiliary views, while physical block placement exposes an instruction-following weakness.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Learning Robot Manipulation from Audio World Models

Fan Zhang; Michael Gienger

An audio world model predicts future spectrograms or MIDI segments, then supplies them to a separately trained robot policy. The paper reports successful physical water filling and qualitative improvement in simulated piano playing, but supplies little comparative numerical evidence. Its key distinction is future audio as an inference-time policy input, with ground-truth futures used during policy training.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video Generation

Guosheng Zhao; Xiaofeng Wang; Zheng Zhu; Xinze Chen; Guan Huang; Xiaoyi Bao; Xingang Wang

DriveDreamer-2 turns language requests into traffic trajectories, trajectory-conditioned road maps and six-view driving videos. UniMVM concatenates camera views so a video diffusion backbone processes them together. Experiments support improved video-distribution scores and synthetic training data for perception; they do not establish closed-loop driving performance.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

A review of learning-based dynamics models for robotic manipulation

Bo Ai; Stephen Tian; Haochen Shi; Yixuan Wang; Tobias Pfaff; Cheston Tan; Henrik I Christensen; Hao Su; Jiajun Wu; Yunzhu Li

This review explains learned dynamics for manipulation through the state that a robot chooses to predict. Pixels, latent vectors, particles, keypoints, and object-centric states redistribute difficulty between perception, prediction, and control. The article connects these choices to trajectory optimization and policy learning, then identifies gaps in partial observability, robustness, and scale. It offers a design taxonomy and literature synthesis rather than a new trained model or comparative benchmark.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

DreamGen: Unlocking Generalization in Robot Learning through Video World Models

Joel Jang; Seonghyeon Ye; Zongyu Lin; Jiannan Xiang; Johan Bjorck; Yu Fang; Fengyuan Hu; Spencer Huang; Kaushil Kundalia; Yen-Chen Lin; Lo\"ıc Magne; Ajay Mandlekar; Avnish Narayan; You Liang Tan; Guanzhi Wang; Jing Wang; Qi Wang; Yinzhen Xu; Xiaohui Zeng; Kaiyuan Zheng; Ruijie Zheng; Ming-Yu Liu; Luke Zettlemoyer; Dieter Fox; Jan Kautz; Scott Reed; Yuke Zhu; Linxi Fan

DreamGen turns embodiment-adapted video generation into offline robot-training data: generate a trajectory, infer its actions, then train a separate visuomotor policy. It improves simulated and real-robot performance, but depends on an action labeler, manual initial frames and substantial generation compute. Reported real-world scores include partial credit. [e02, e04, e09, e11, e13, e16]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Latent Action Pretraining from Videos

Seonghyeon Ye; Joel Jang; Byeongguk Jeon; Sejune Joo; Jianwei Yang; Baolin Peng; Ajay Mandlekar; Reuben Tan; Yu-Wei Chao; Bill Yuchen Lin; Lars Liden; Kimin Lee; Jianfeng Gao; Luke Zettlemoyer; Dieter Fox; Minjoon Seo

LAPA turns visual changes into discrete training targets for a vision-language-action policy, then uses labeled robot demonstrations to learn executable actions. It improves transfer from actionless videos, including human videos, while retaining weaknesses in precise grasping. Its separate decoder supports qualitative imagined rollouts; the robot results evaluate the fine-tuned policy (e02, e05, e12, e13, e14, e20).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Precise and Dexterous Robotic Manipulation via Human-in-the-Loop Reinforcement Learning

Jianlan Luo; Charles Xu; Jeffrey Wu; Sergey Levine

HIL-SERL combines visual model-free reinforcement learning, demonstrations, corrective teleoperation and task-specific robot control. Its central mechanism uses human recoveries as replayable transitions, allowing reward-based learning to improve beyond imitation. The paper reports perfect observed success across its evaluation tasks, usually after 1–2.5 hours of online training, with a six-hour timing-belt exception. These are physical task results under controlled setups, not demonstrations of a predictive world model or generalist deployment.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Mean Flows for One-step Generative Modeling

Zhengyang Geng; Mingyang Deng; Xingjian Bai; Zico Kolter; Kaiming He

MeanFlow learns the average velocity between two times, allowing a noise sample to cross an entire generative flow in one network evaluation. A differential identity supplies a training target without integrating trajectories or distilling a teacher. Classifier-free guidance is learned into the field. ImageNet results are strong, but the pretrained latent tokenizer, single-run evaluation and configuration changes qualify the headline claims (e03–e06, e08, e11, e14, e21).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Helix: A Vision-Language-Action Model for Generalist Humanoid Control

Figure AI

Helix connects an internet-pretrained vision-language model to a faster visuomotor policy through a learned continuous latent vector. Figure's overview describes natural-language-conditioned upper-body control, including novel-object grasping and two-robot grocery storage. S2 supplies semantic updates at 7–9 Hz; S1's control loop outputs at 200 Hz, while the original architecture labels S1's image/state input branches at 20 Hz. Joint action-regression training includes a latency-matched input offset. The inspected figures explain the motivation and architecture; task outcomes remain publisher-reported accounts without success-rate denominators, controlled ablations or measured scaling curves.

Resource reviewedIllustrated editionFull text availableRead the report Primary source

AdaWorld: Learning Adaptable World Models with Latent Actions

Shenyuan Gao; Siyuan Zhou; Yilun Du; Jun Zhang; Chuang Gan

AdaWorld learns a continuous description of change between video frames, then uses that description to condition a separate next-frame diffusion model. Demonstration actions can be replayed in a new visual context, or mapped to an environment's controls with limited adaptation. Its central contribution is pretraining an adaptable control interface; planning still requires an external optimizer and environment feedback.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models

Weixin Liang; Lili Yu; Liang Luo; Srinivasan Iyer; Ning Dong; Chunting Zhou; Gargi Ghosh; Mike Lewis; Wen-tau Yih; Luke Zettlemoyer; Xi Victoria Lin

Mixture-of-Transformers (MoT) assigns modality-specific weights throughout transformer layers while exchanging information through global attention. It accelerates multimodal pretraining, especially image and speech modeling; benefits depend on the objective and evaluation protocol. This is a generative backbone study, with no action policy or executed-control evaluation. [e-routing, e-chameleon-result, e-speech, e-text-boundary]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

FAST-LIVO2: Fast, Direct LiDAR-Inertial-Visual Odometry

Chunran Zheng; Wei Xu; Zuhao Zou; Tong Hua; Chongjian Yuan; Dongjiao He; Bingyang Zhou; Zheng Liu; Jiarong Lin; Fangcheng Zhu; Yunfan Ren; Rong Wang; Fanle Meng; Fu Zhang

FAST-LIVO2 estimates sensor motion and builds a colored local map by combining IMU propagation, direct LiDAR registration and sparse image-patch alignment. Its central design is a shared voxel map and a LiDAR-first, camera-second filter update. The reported gains concern odometry and mapping; autonomous flight additionally uses a separate planner and controller.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

RLVR-World: Training World Models with Reinforcement Learning

Jialong Wu; Shaofeng Yin; Ningya Feng; Mingsheng Long

RLVR-World post-trains autoregressive world models by scoring decoded next-state predictions against ground truth and reinforcing better samples. Separate language and video implementations improve transition metrics and selected downstream applications, while remaining dependent on the base model, reward and evaluation protocol (e02–e05, e09–e14).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

ParticleFormer: A 3D Point Cloud World Model for Multi-Object, Multi-Material Robotic Manipulation

Suning Huang; Qianzhong Chen; Xiaohan Zhang; Jiankai Sun; Mac Schwager

ParticleFormer predicts action-conditioned 3D particle motion using material-aware Transformer attention and a Chamfer–Hausdorff training loss. Stereo reconstruction and segmentation supply point sets without requiring tracked particle correspondences. A separate MPPI planner uses predicted futures for manipulation. Experiments support better MSE and combined geometric error than the tested baselines, with a Chamfer-only tradeoff; deployment remains scene-specific. [e03, e05, e06, e07, e09, e13]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

π₀.₅: a Vision-Language-Action Model with Open-World Generalization

Kevin Black; Noah Brown; James Darpinian; Karan Dhabalia; Danny Driess; Adnan Esmail; Michael Robert Equi; Chelsea Finn; Niccolo Fusai; Manuel Y. Galliker; Dibya Ghosh; Lachy Groom; Karol Hausman; brian ichter; Szymon Jakubczak; Tim Jones; Liyiming Ke; Devin LeBlanc; Sergey Levine; Adrian Li-Bell; Mohith Mothukuri; Suraj Nair; Karl Pertsch; Allen Z. Ren; Lucy Xiaoyang Shi; Laura Smith; Jost Tobias Springenberg; Kyle Stachowicz; James Tanner; Quan Vuong; Homer Walke; Anna Walling; Haohuan Wang; Lili Yu; Ury Zhilinsky

π0.5 transfers heterogeneous robot, semantic and web supervision into a shared vision-language-action policy for household manipulation in unseen homes. Discrete action tokens support pre-training; a flow-matching expert supplies continuous control after post-training. The experiments demonstrate executed robot task progress, while the hierarchy ablation separates training benefits from the less certain incremental benefit of explicit subtask inference (e02, e04, e05, e11, e17).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

DrivingGPT: Unifying Driving World Modeling and Planning with Multi-modal Autoregressive Transformers

Yuntao Chen; Yuqi Wang; Zhaoxiang Zhang

DrivingGPT turns front-camera frames and relative ego motion into interleaved discrete tokens. A single causal transformer predicts both future imagery and motion, then integrates motion into trajectories. Its strongest planning comparison is on NAVSIM navmini, with separate navtest video-quality evaluations. The results support a unified prediction architecture, while leaving reactive driving, causal benefits of joint training and deployment efficiency unresolved.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Mastering diverse control tasks through world models

Danijar Hafner; Jurgis Pasukonis; Jimmy Ba; Timothy Lillicrap

DreamerV3 learns latent dynamics from experience and trains an actor through imagined trajectories. Its contribution is robust learning across diverse control domains using shared algorithm settings, with benchmark-specific compute choices. Minecraft diamond discovery is demonstrated under a modified MineRL interface. Evidence: e-world, e-actor, e-protocol-table, e-minecraft-setup, e-minecraft.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

World Model-based Perception for Visual Legged Locomotion

Hang Lai; Jiahang Cao; Jiafeng Xu; Hongtao Wu; Yunfeng Lin; Tao Kong; Yong Yu; Weinan Zhang

World Model-based Perception (WMP) trains a recurrent sensor-prediction model alongside a separate locomotion policy in simulation. The model supplies remembered context to the controller; PPO uses simulator trajectories rather than imagined rollouts. Simulation comparisons and Unitree A1 trials support improved terrain traversal, while the hardest physical tests still have substantial failure rates.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning

Gaoyue Zhou; Hengkai Pan; Yann LeCun; Lerrel Pinto

DINO-WM learns how actions change frozen DINOv2 patch features, then searches for actions that bring predicted features close to a goal image. Its strongest evidence concerns simulated manipulation under an offline, reward-free planning protocol. Spatial features and causal prediction matter, but data provenance, appendix discrepancies and planning latency constrain the zero-shot claim.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Evaluating Real-World Robot Manipulation Policies in Simulation

Xuanlin Li; Kyle Hsu; Jiayuan Gu; Oier Mees; Karl Pertsch; Homer Rich Walke; Chuyuan Fu; Ishikaa Lunawat; Isabel Sieh; Sean Kirmani; Sergey Levine; Jiajun Wu; Chelsea Finn; Hao Su; Quan Vuong; Ted Xiao

SIMPLER builds purpose-specific physics simulations to evaluate real-data-trained manipulation policies. It calibrates action execution and visual observations, then tests whether simulated successes preserve real policy rankings and robustness trends. Its strongest evidence concerns rigid-object tasks on Google Robot and WidowX; useful ranking signals coexist with absolute-success errors and reporting inconsistencies (e02, e10–e13, e23).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Epona: Autoregressive Diffusion World Model for Autonomous Driving

Kaiwen Zhang; Zhenyu Tang; Xiaotao Hu; Xingang Pan; Xiaoyang Guo; Yuan Liu; Jingwei Huang; Li Yuan; Qian Zhang; Xiao-Xiao Long; Xun Cao; Wei Yin

Epona compresses camera and ego-motion history into a shared temporal representation, then uses separate diffusion transformers to predict a continuous future trajectory and one action-conditioned image. Repeating image prediction produces long videos; disabling it enables cheaper trajectory inference. Its strongest mechanism evidence is improved planning after joint visual training and reduced rollout degradation after Chain-of-Forward training. Reported video quality, qualitative duration and benchmark planning scores establish different capabilities, with important protocol and implementation gaps.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Robotic World Model: A Neural Network Simulator for Robust Policy Optimization in Robotics

Chenhao Li; Andreas Krause; Marco Hutter

RWM trains a recurrent neural simulator on its own multi-step predictions, then uses it to train a separate PPO controller. The paper demonstrates trajectory prediction and zero-shot locomotion-policy transfer, while retaining simulation pretraining and simulation-based online data collection. Its central tradeoff is better long-rollout fidelity at greater world-model training cost [e-ar, e-policy, e-pretraining, e-hardware, e-ablation].

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

EgoBridge: Domain Adaptation for Generalizable Imitation from Egocentric Human Data

Ryan Punamiya; Dhruv Patel; Patcharapong Aphiwetsa; Pranav Kuppili; Lawrence Zhu; Simar Kareer; Judy Hoffman; Danfei Xu

EgoBridge co-trains robot policies with labeled egocentric human demonstrations, using motion similarity to guide optimal-transport alignment of observation features. Its contribution is training-time representation transfer into executed manipulation policies. Evidence includes transfer to robot-unseen drawer locations and visual settings, with unresolved implementation details and unreported uncertainty.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

World4Drive: End-to-End Autonomous Driving via Intention-aware Physical Latent World Model

Yupeng Zheng; Pengxuan Yang; Zebin Xing; Qichao Zhang; Yuhang Zheng; Yinfeng Gao; Pengfei Li; Teng Zhang; Zhongpu Xia; Peng Jia; XianPeng Lang; Dongbin Zhao

World4Drive proposes several command-conditioned driving trajectories, predicts a future latent scene for each, and learns to rank them by resemblance to an observed future. Depth-derived positions, semantic pseudo-labels and temporal attention enrich its scene representation. It improves reported average collision metrics over perception-free LAW, while retaining expert-trajectory supervision and exhibiting important accuracy–collision tradeoffs.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

A Survey on Vision-Language-Action Models for Autonomous Driving

Sicong Jiang; Zilin Huang; Kangan Qian; Ziang Luo; Tianze Zhu; Yang Zhong; Yihong Tang; Menglin Kong; Yunlong Wang; Siwen Jiao; Hao Ye; Zihao Sheng; Xin Zhao; Tuopu Wen; Zheng Fu; Sikai Chen; Kun Jiang; Diange Yang; Seongjin Choi; Lijun Sun

This survey organizes autonomous-driving VLA research around how language enters the path from sensing to vehicle action. It distinguishes explanatory overlays, modular language-to-plan interfaces, unified policies and reasoning augmentation. Its contribution is a vocabulary for comparing inputs, representations, action outputs, training and evaluation. It supplies conceptual diagrams and literature inventories rather than a new trained controller or controlled benchmark ranking. Several descriptions conflict across prose, tables and references, so individual model claims require cautious attribution. The architectural and evaluation synthesis is supported by e-gap, e-architecture, e-stages and e-evaluation.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-Tuning

Zewei Zhou; Tianhui Cai; Seth Zhao; Yun Zhang; Zhiyu Huang; Bolei Zhou; Jiaqi Ma

AutoVLA extends Qwen2.5-VL-3B with discrete vehicle-motion tokens so one autoregressive decoder can produce scene reasoning and a driving trajectory. Supervised training teaches short and extended responses; GRPO then rewards planning quality while penalizing long reasoning. Gains depend on data, decoding and evaluation protocol, and physical execution is demonstrated in simulation.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation

Yang Tian; Sizhe Yang; Jia Zeng; Ping Wang; Dahua Lin; Hao Dong; Jiangmiao Pang

Seer learns manipulation by forecasting a future visual latent and conditioning an inverse-dynamics action chunk on that latent inside one transformer. Joint robot-data pre-training improves subsequent task learning, but the strongest CALVIN result belongs to Seer-Large, and transfer from other embodiments is much weaker than DROID transfer. The evidence below separates predictive representations, executed-task metrics, and unresolved reporting details.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

π₀: A Vision-Language-Action Flow Model for General Robot Control

Kevin Black; Noah Brown; Danny Driess; Adnan Esmail; Michael Robert Equi; Chelsea Finn; Niccolo Fusai; Lachy Groom; Karol Hausman; Brian Ichter; Szymon Jakubczak; Tim Jones; Liyiming Ke; Sergey Levine; Adrian Li-Bell; Mohith Mothukuri; Suraj Nair; Karl Pertsch; Lucy Xiaoyang Shi; Laura Smith; James Tanner; Quan Vuong; Anna Walling; Haohuan Wang; Ury Zhilinsky

π0 adapts a pre-trained vision-language model to continuous robot control through a smaller flow-matching action expert. Diverse robot pre-training supplies a reusable policy, while task-specific post-training improves dexterous execution. Real-robot experiments support this recipe, but performance depends on task and training conditions; partial-credit scores and several source inconsistencies limit broad claims of mastery.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins

Yao Mu; Tianxing Chen; Zanxin Chen; Shijia Peng; Zhiqian Lan; Zeyu Gao; Zhixuan Liang; Qiaojun Yu; Yude Zou; Mingkun Xu; Lunkai Lin; Zhiqiang Xie; Mingyu Ding; Ping Luo

RoboTwin turns object photographs into annotated simulation assets and uses language-generated programs plus motion planning to collect expert demonstrations. It benchmarks imitation policies and tests simulation pretraining on physical robots. The strongest evidence is improved performance with limited real demonstrations; difficult dual-arm coordination remains unresolved. [e-problem, e-code, e-dual, e-boundaries]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

FMB: A functional manipulation benchmark for generalizable robotic learning

Jianlan Luo; Charles Xu; Fangchen Liu; Liam Tan; Zipeng Lin; Jeffrey Wu; Pieter Abbeel; Sergey Levine

FMB turns assembly into a controlled test of functional grasping, reorientation, insertion and skill composition. Its physical object set and segmented demonstrations support modular imitation-learning baselines. The central finding is that human-oracle skill sequencing enables some complete assemblies where the tested flat policies fail; autonomous high-level planning remains untested (e02, e05, e10, e16, e17).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Zhuoyi Yang; Jiayan Teng; Wendi Zheng; Ming Ding; Shiyu Huang; Jiazheng Xu; Yuanming Yang; Wenyi Hong; Xiaohan Zhang; Guanyu Feng; Da Yin; Yuxuan Zhang; Weihan Wang; Yean Cheng; Bin Xu; Xiaotao Gu; Yuxiao Dong; Jie Tang

CogVideoX generates videos from text using a temporally compressing VAE and a diffusion transformer that jointly processes text and visual tokens. Its distinguishing choice is modality-specific adaptive normalization around shared attention and feed-forward computation. Full spatiotemporal attention supports information exchange across moving objects, while frame packing, progressive resolution training and dense captions address data efficiency and alignment. The inspected revision reports generation up to 10 seconds at 16 fps and 768 × 1360 pixels. Experiments support several design choices and competitive video quality, with important protocol gaps and internal reporting discrepancies. The output is video; no action policy or environment-control loop is provided.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

A0: An Affordance-Aware Hierarchical Model for General Robotic Manipulation

Rongtao Xu; Jian Zhang; Minghao Guo; Youpeng Wen; Haoting Yang; Min Lin; Jianzheng Huang; Zhe Li; Kaidong Zhang; Liqiong Wang; Yuxuan Kuang; Meng Cao; Feng Zheng; Xiaodan Liang

A0 predicts an object contact point and a short 2D post-contact trajectory from an image and instruction. A separate depth, grasp and motion pipeline executes those points. Its strongest evidence combines lower offline waypoint error with real Franka and Kinova task success; embodiment independence remains conditional on the execution stack and tested settings.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Diffusion policy: Visuomotor policy learning via action diffusion

Cheng Chi; Zhenjia Xu; Siyuan Feng; Eric Cousineau; Yilun Du; Benjamin Burchfiel; Russ Tedrake; Shuran Song

Diffusion Policy learns a conditional distribution over demonstrated action sequences, then repeatedly denoises a candidate sequence and executes a short segment before observing again. Visual features condition action generation rather than being predicted as future images. The journal extension adds encoder studies and bimanual manipulation. Its strongest evidence combines simulation comparisons with executed real-robot tasks, while heterogeneous protocols and conflicting configuration statements limit simple headline comparisons.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Pseudo-Simulation for Autonomous Driving

Wei Cao; Marcel Hallgarten; Tianyu Li; Daniel Dauner; Xunjiang Gu; Caojun Wang; Yakov Miron; Marco Aiello; Hongyang Li; Igor Gilitschenski; Boris Ivanovic; Marco Pavone; Andreas Geiger; Kashyap Chitta

Pseudo-simulation evaluates driving planners on recorded observations and a fixed bank of synthetic future observations. Gaussian proximity weights connect the two stages without online sensor rendering. NAVSIM v2 thereby tests responses to deviations from expert driving while permitting parallel evaluation. Its strongest evidence is improved correlation with nuPlan closed-loop scores; it does not establish real-world safety or a learned world-action architecture (e-flow, e-correlation, e-limitations).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations

Yucheng Hu; Yanjiang Guo; Pengchao Wang; Xiaoyu Chen; Yen-Jen Wang; Jianke Zhang; Koushil Sreenath; Chaochao Lu; Jianyu Chen

Video Prediction Policy (VPP) turns a manipulation-adapted video diffusion model into a predictive visual encoder. A separate diffusion policy learns actions from its internal future features. One video forward pass replaces full video generation during control. CALVIN and physical-robot results support the approach, while conflicting values and incomplete protocol details constrain causal and reproducibility conclusions.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

DyWA: Dynamics-adaptive World Action Model for Generalizable Non-prehensile Manipulation

Jiangran Lyu; Ziming Li; Xuesong Shi; Chaoyi Xu; Yizhou Wang; He Wang

DyWA learns contact-rich object rearrangement from a single-view partial point cloud. A privileged simulation teacher supervises a student that jointly predicts robot actions and the next object-to-goal transformation. Historical observations and actions supply a dynamics embedding that modulates the student through FiLM. The strongest evidence combines controlled simulation ablations with a small physical-robot evaluation; the method does not demonstrate planning through imagined rollouts.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Navigation World Models

Amir Bar; Gaoyue Zhou; Danny Tran; Trevor Darrell; Yann LeCun

Navigation World Models turns action-conditioned video prediction into a navigation evaluator: simulate candidate trajectories, compare their final views with a goal image, and select actions through search or policy reranking. The strongest main-paper result concerns short RECON trajectory errors; unfamiliar-scene imagination remains vulnerable to drift.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Cross-Embodiment Dexterous Manipulation through World Model Learning

Zihao He; Bo Ai; Yulin Liu; Weikang Wan; Henrik I Christensen; Hao Su

This revised paper shares a particle-based dynamics predictor across human and robot hands, then uses each robot’s forward kinematics and model-predictive control to choose feasible actions. Simulation studies examine embodiment diversity; real clay-reshaping experiments compare human-only and mixed-domain training. The strongest quantitative evidence is lower final shape error with co-training on two physical hands. Generalization remains conditional on perception, known kinematics and restricted action primitives.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

GWM: Towards Scalable Gaussian World Models for Robotic Manipulation

Guanxing Lu; Baoxiong Jia; Puhao Li; Yixin Chen; Ziwei Wang; Yansong Tang; Siyuan Huang

GWM lifts RGB observations into Gaussian splats, compresses them, and predicts action-conditioned future scene latents with diffusion. Separate policies use these features for imitation learning or learn from imagined transitions. The experiments support improved average manipulation performance, with task regressions and unresolved implementation discrepancies that qualify the scalability claims.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

X-MOBILITY: End-To-End Generalizable Navigation via World Modeling

Wei Liu; Huihua Zhao; Chenran Li; Joydeep Biswas; Billy Okal; Pulkit Goyal; Yan Chang; Soha Pouya

X-MOBILITY learns recurrent navigation beliefs from synthetic random-action experience, then couples them to route-conditioned imitation. Semantic and RGB decoding shape the representation; a separate predictor forecasts latent consequences without new observations. Stronger closed-loop simulation results coexist with modest obstacle-prediction accuracy and limited real-robot evidence (e03, e05, e06, e08, e12, e14, e15, e16).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

OpenVLA: An Open-Source Vision-Language-Action Model

Moo Jin Kim; Karl Pertsch; Siddharth Karamcheti; Ted Xiao; Ashwin Balakrishna; Suraj Nair; Rafael Rafailov; Ethan Paul Foster; Pannag R. Sanketi; Quan Vuong; Thomas Kollar; Benjamin Burchfiel; Russ Tedrake; Dorsa Sadigh; Sergey Levine; Percy Liang; Chelsea Finn

OpenVLA turns a pretrained vision-language model into a direct robot policy by predicting discretized action tokens from one image and an instruction. Its practical strengths are generalization and adaptation across manipulation settings. The evidence also exposes important boundaries: some scores include partial credit, efficiency tests use a smaller variant, and control latency changes apparent quantization quality.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Driving in the Occupancy World: Vision-Centric 4D Occupancy Forecasting and Planning via World Models for Autonomous Driving

Yu Yang; Jianbiao Mei; Yukai Ma; Siliang Du; Wenqing Chen; Yijie Qian; Yuxiang Feng; Yong Liu

Drive-OccWorld turns historical camera images into future semantic occupancy and flow, then uses those predictions to score and refine ego trajectories. Its central contribution is an inference-time feedback loop between a BEV world model and an occupancy-based planner. Reported nuScenes improvements concern forecasting and open-loop planning; they do not establish deployed driving safety.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

ManiGaussian++: General Robotic Bimanual Manipulation with Hierarchical Gaussian World Model

Tengbo Yu; Guanxing Lu; Zaijia Yang; Haoyuan Deng; Season Si Chen; Jiwen Lu; Wenbo Ding; Guoqiang Hu; Yansong Tang; Ziwei Wang

ManiGaussian++ improves language-conditioned bimanual imitation learning by training visual features to reconstruct task-labeled Gaussian scenes and predict their action-conditioned deformation. A stabilizing-arm leader precedes an acting-arm follower; a separate policy maps the enriched features to robot actions. The strongest evidence is higher multi-task success, with unresolved protocol details and a discrepancy between headline and detailed real-world averages.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems

AgiBot-World-Contributors; Qingwen Bu; Jisong Cai; Li Chen; Xiuqi Cui; Yan Ding; Siyuan Feng; Shenyuan Gao; Xindong He; Xuan Hu; Xu Huang; Shu Jiang; Yuxin Jiang; Cheng Jing; Hongyang Li; Jialu Li; Chiming Liu; Yi Liu; Yuxiang Lu; Jianlan Luo; Ping Luo; Yao Mu; Yuehan Niu; Yixuan Pan; Jiangmiao Pang; Yu Qiao; Guanghui Ren; Cheng Ruan; Jiaqi Shan; Yongjian Shen; Chengshi Shi; Mingkang Shi; Modi Shi; Chonghao Sima; Jianheng Song; Huijie Wang; Wenhao Wang; Dafeng Wei; Chengen Xie; Guo Xu; Junchi Yan; Cunbiao Yang; Lei Yang; Shukai Yang; Maoqing Yao; Jia Zeng; Chi Zhang; Qinglin Zhang; Bin Zhao; Chengyue Zhao; Jiaqi Zhao; Jianchao Zhu

AgiBot World Colosseo combines a million-trajectory manipulation dataset, standardized collection and the GO-1 policy. GO-1 learns discrete motion representations from videos, predicts those representations from images and language, and uses them to condition continuous robot controls. Real-world experiments support dataset-transfer and policy gains on selected tasks; they do not establish general competence across the entire collection.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Inference-Time Enhancement of Generative Robot Policies via Predictive World Modeling

Han Qi; Haocheng Yin; Aris Zhu; Yilun Du; Heng Yang

Generative Predictive Control (GPC) improves a frozen diffusion policy by predicting the consequences of its action proposals, then selecting or refining them before execution. Its separate world model learns from demonstrations plus random exploration. Simulation and small real-world studies support this combination, while rollout cost and incomplete evaluation details limit the strength of broader deployment claims (e02, e03, e10, e12, e16, e18).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Enhancing End-to-End Autonomous Driving with Latent World Model

Yingyan Li; Lue Fan; Jiawei He; Yuqi Wang; Yuntao Chen; Zhaoxiang Zhang; Tieniu Tan

LAW trains an end-to-end driving planner to predict future visual features from current features and its predicted ego trajectory. This auxiliary task works with perspective-view or BEV representations. Controlled ablations, open-loop driving metrics and CARLA closed-loop scores support useful representation learning, while leaving counterfactual dynamics accuracy and physical deployment untested.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Doe-1: Closed-Loop Autonomous Driving with Large World Model

Wenzhao Zheng; Zetian Xia; Yuanhui Huang; Sicheng Zuo; Jie Zhou; Jiwen Lu

Doe-1 puts front-camera images, free-form scene descriptions and discretized ego motion into one autoregressive token stream. Different prompts expose perception, action-conditioned image prediction and trajectory planning. Its nuScenes experiments support this shared capability, while the claimed closed loop is demonstrated through model-generated scene evolution rather than a quantitative driving evaluation with live environment feedback.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Chi-Lam Cheang; Guangzeng Chen; Ya Jing; Tao Kong; Hang Li; Yifeng Li; Yuxiao Liu; Hongtao Wu; Jiafeng Xu; Yichu Yang; Hanbo Zhang; Minzhao Zhu

GR-2 learns language-conditioned video prediction at scale, then adapts a GPT-style policy to predict robot action trajectories alongside future images. Physical multi-task and bin-picking results and simulated CALVIN chains support its practical capability, while qualitative video alignment and model scaling leave the causal contribution of visual planning unresolved.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Scaling Offline Model-Based RL via Jointly-Optimized World-Action Model Pretraining

Jie Cheng; Ruixi Qiao; Yingwei Ma; Binhua Li; Gang Xiong; Qinghai Miao; Yongbin Li; Yisheng Lv

JOWA learns an Atari world model and distributional Q-function through one shared transformer, then searches short imagined futures to choose actions. Its strongest evidence is improved aggregate game return and data-efficient offline adaptation; neither universal game-wise scaling nor unconditional planning optimality is established (e02, e03, e04, e07, e08, e11, e12).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Homanga Bharadhwaj; Debidatta Dwibedi; Abhinav Gupta; Shubham Tulsiani; Carl Doersch; Ted Xiao; Dhruv Shah; Fei Xia; Dorsa Sadigh; Sean Kirmani

Gen2Act turns a language instruction and an initial scene image into a generated human demonstration, then uses that video to condition a separate closed-loop robot policy. A pretrained VideoPoet supplies the demonstration without robot-specific fine-tuning. Auxiliary point-track prediction teaches policy representations to retain motion cues, while deployment predicts actions directly from video features and recent robot observations. Real robot results support improved generalization relative to the reported baselines, but plausible generation does not ensure correct execution, and long-horizon reliability remains limited.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

DynaMo: In-Domain Dynamics Pretraining for Visuo-Motor Control

Zichen Jeff Cui; Hengkai Pan; Aadhithya Iyer; Siddhant Haldar; Lerrel Pinto

DynaMo uses action-free dynamics pretraining to make visual features useful for imitation learning. An encoder, latent inverse model and forward model jointly learn to predict next-frame embeddings; a separate policy then learns from frozen features and labeled demonstrations. Results favor DynaMo on several manipulation tasks, but include ties and initialization regressions. The contribution is a representation-learning objective, with no demonstrated use of the dynamics models for online planning.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

OccLLaMA: An Occupancy-Language-Action Generative World Model for Autonomous Driving

Julong Wei; Shanshuai Yuan; Pengfei Li; Qingda Hu; Zhongxue Gan; Wenchao Ding

OccLLaMA makes semantic 3D occupancy, language, and ego waypoints share a generative vocabulary. A sparse scene tokenizer feeds a modified LLaMA that predicts text/actions sequentially and all tokens within each occupancy scene in parallel. NuScenes-based results support multitask forecasting, planning, and question answering, while the component ablation exposes a forecasting–planning tradeoff.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

PaliGemma: A versatile 3B VLM for transfer

Lucas Beyer; Andreas Steiner; André Susano Pinto; Alexander Kolesnikov; Xiao Wang; Daniel Salz; Maxim Neumann; Ibrahim Alabdulmohsin; Michael Tschannen; Emanuele Bugliarello; Thomas Unterthiner; Daniel Keysers; Skanda Koppula; Fangyu Liu; Adam Grycner; Alexey Gritsenko; Neil Houlsby; Manoj Kumar; Keran Rong; Julian Eisenschlos; Rishabh Kabra; Matthias Bauer; Matko Bošnjak; Xi Chen; Matthias Minderer; Paul Voigtlaender; Ioana Bica; Ivana Balazevic; Joan Puigcerver; Pinelopi Papalampidi; Olivier Henaff; Xi Xiong; Radu Soricut; Jeremiah Harmsen; Xiaohua Zhai

PaliGemma connects a SigLIP vision encoder to Gemma-2B, then trains the entire VLM to become a useful starting point for task-specific adaptation. Its distinctive choices are bidirectional image–prompt attention, suffix-only prediction loss, broad multimodal pretraining and separate resolution checkpoints. Strong results concern visual understanding and structured visual outputs; the paper does not demonstrate action execution or future-world prediction.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

This&That: Language-Gesture Controlled Video Generation for Robot Planning

Boyang Wang; Nikhil Sridhar; Chao Feng; Mark Van der Merwe; Adam Fishman; Nima Fazeli; Jeong Joon Park

This&That turns an initial image, language and pointing coordinates into a generated video plan, then uses a separate DiVA behavior-cloning controller to follow it with live feedback. Gestures improve intent disambiguation in Bridge video generation and simulated block manipulation; real-robot execution remains untested.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Dreamitate: Real-World Visuomotor Policy Learning via Video Generation

Junbang Liang; Ruoshi Liu; Ege Ozguroglu; Sruthi Sudhakar; Achal Dave; Pavel Tokmakov; Shuran Song; Carl Vondrick

Dreamitate learns task-specific video generation from human tool demonstrations, then extracts tool poses from synthesized stereo videos for robot execution. It outperforms the tested Diffusion Policy baseline on four physical manipulation tasks, while depending on calibrated cameras, known rigid tools, and slow open-loop generation. The central evidence concerns executed tool trajectories under specified object and scene shifts, rather than unrestricted manipulation or a general action-conditioned simulator.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

RoboDreamer: Learning Compositional World Models for Robot Imagination

Siyuan Zhou; Yilun Du; Jiaben Chen; Yandong Li; Dit-Yan Yeung; Chuang Gan

RoboDreamer composes phrase-conditioned video diffusion predictions to imagine robot plans for unfamiliar instruction combinations. Optional goal images or sketches sharpen spatial specifications; a separate inverse-dynamics model converts imagined frames into actions. The strongest language-only evidence concerns human-rated video alignment, while executed success is measured separately in RLBench simulation (E03, E10–E13).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Prediction with Action: Visual Policy Learning via Joint Denoising Process

Yanjiang Guo; Yucheng Hu; Jianke Zhang; Yen-Jen Wang; Xiaoyu Chen; Chaochao Lu; Jianyu Chen

PAD learns a language-conditioned robot policy by jointly denoising future image latents and robot poses in one diffusion transformer. RGB video training supplies prediction experience without requiring action labels. The robot executes the first predicted pose and observes again. Reported control gains are substantial, but the MetaWorld headline needs qualification because the appendix excludes two tasks.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

NAVSIM: Data-Driven Non-Reactive Autonomous Vehicle Simulation and Benchmarking

Daniel Dauner; Marcel Hallgarten; Tianyu Li; Xinshuo Weng; Zhiyu Huang; Zetong Yang; Hongyang Li; Igor Gilitschenski; Boris Ivanovic; Marco Pavone; Andreas Geiger; Kashyap Chitta

NAVSIM benchmarks sensor-based driving by simulating a fixed four-second plan against non-reactive traffic, then scoring safety, progress and comfort. Its curated scenes expose the weakness of blind ego-state policies, while compact sensor models match more elaborate baselines. The evidence supports an evaluation proxy, not a guarantee of reactive driving competence. [e03, e05, e08, e13]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation

Hongtao Wu; Ya Jing; Chilam Cheang; Guangzeng Chen; Jiafeng Xu; Xinghang Li; Minghuan Liu; Hang Li; Tao Kong

GR-1 transfers language-conditioned future-frame prediction on human videos into a robot policy. A shared causal transformer predicts images and actions after robot fine-tuning. The experiments show better CALVIN task-chain completion and real-robot manipulation than the tested baselines, while ablations support contributions from both video supervision and pretraining. These are executed-policy results; generated images alone do not establish control success.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

DINOv2: Learning Robust Visual Features without Supervision

Maxime Oquab; Timothée Darcet; Théo Moutakanni; Huy V. Vo; Marc Szafraniec; Vasil Khalidov; Pierre Fernandez; Daniel Haziza; Francisco Massa; Alaaeldin El-Nouby; Mido Assran; Nicolas Ballas; Wojciech Galuba; Russell Howes; Po-Yao Huang; Shang-Wen Li; Ishan Misra; Michael Rabbat; Vasu Sharma; Gabriel Synnaeve; Hu Xu; Hervé Jégou; Julien Mairal; Patrick Labatut; Armand Joulin; Piotr Bojanowski

DINOv2 learns reusable image and patch representations through curated image data, DINO/iBOT self-supervision, efficient large-model training and distillation. Its strongest evidence concerns frozen visual features used by task-specific heads: ImageNet linear accuracy reaches 86.5%, while dense prediction and retrieval expose complementary benefits of the patch objective and feature-spreading regularizer. These are perception results, with no demonstrated action-conditioned dynamics or control loop.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots

Cheng Chi; Zhenjia Xu; Chuer Pan; Eric Cousineau; Benjamin Burchfiel; Siyuan Feng; Russ Tedrake; Shuran Song

UMI turns handheld gripper demonstrations into deployable visuomotor policies by matching camera geometry, recovering metric actions and aligning robot timing. Its contribution is a collection-and-control interface with real robot evaluations, rather than a learned predictor of future world observations.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

World Models via Policy-Guided Trajectory Diffusion

Marc Rigter; Jun Yamada; Ingmar Posner

PolyGRAD generates approximately on-policy synthetic state, reward, and action trajectories by combining an action-conditioned denoiser with a separate policy's action score. It refines the whole trajectory through diffusion instead of rolling out one transition at a time. Short-horizon prediction is competitive, but imagined-policy learning remains weaker than Dreamer-v3, and low-entropy action calibration is fragile.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Bench2Drive: Towards Multi-Ability Benchmarking of Closed-Loop End-To-End Autonomous Driving

Xiaosong Jia; Zhenjie Yang; Qifeng Li; Zhiyuan Zhang; Junchi Yan

Bench2Drive standardizes demonstrations and tests end-to-end driving through short, interactive CARLA routes. Its contribution is an evaluation system: a common expert dataset, scenario-level success measurements and five ability groups. Baselines reveal gaps between fitting recorded trajectories and successful closed-loop driving. Expert-feature distillation improves aggregate TCP-traj performance, but skill-specific regressions and simulator limitations constrain the interpretation.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Generalized Predictive Model for Autonomous Driving

Jiazhi Yang; Shenyuan Gao; Yihang Qiu; Li Chen; Tianyu Li; Bo Dai; Kashyap Chitta; Penghao Wu; Jia Zeng; Ping Luo; Jun Zhang; Andreas Geiger; Yu Qiao; Hongyang Li

GenAD adapts an image diffusion model into a driving-video predictor using OpenDV-2K and trainable temporal reasoning blocks. Its future-video generator can receive trajectory conditions, while a separate lightweight planner reads frozen encoder features. The evidence supports visual prediction and economical planning adaptation, with substantial pretraining and limited deployment validation. [e-data-scale, e-temporal, e-extensions, e-plan, e-limits]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

State Estimation for Robotics

Timothy D. Barfoot

This selected-chapter reading follows how noisy measurements become state and pose estimates with uncertainty. Chapter 3 connects Gaussian least squares, exact smoothing, causal filtering and a special sparse continuous-time GP construction. Chapter 8 extends estimation to rigid-body geometry through alignment, tracking and pose-graph relaxation. The common lesson is to respect information structure and geometric constraints. These are foundational estimation tools, not a learned action-generating system. The reviewed source is the corrected first edition compiled in 2022, not the cataloged 2024 second edition.

Partially reviewedIllustrated editionPartial text availableRead the report Primary source

DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

Alexander Khazatsky; Karl Pertsch; Suraj Nair; Ashwin Balakrishna; Sudeep Dasari; Siddharth Karamcheti; Soroush Nasiriany; Mohan Kumar Srirama; Lawrence Yunliang Chen; Kirsty Ellis; Peter David Fagan; Joey Hejna; Masha Itkina; Marion Lepert; Yecheng Jason Ma; Patrick Tree Miller; Jimmy Wu; Suneel Belkhale; Shivin Dass; Huy Ha; Arhan Jain; Abraham Lee; Youngwoon Lee; Marius Memmel; Sungjae Park; Ilija Radosavovic; Kaiyuan Wang; Albert Zhan; Kevin Black; Cheng Chi; Kyle Beltran Hatch; Shan Lin; Jingpei Lu; Jean Mercat; Abdul Rehman; Pannag R. Sanketi; Archit Sharma; Cody Simpson; Quan Vuong; Homer Rich Walke; Blake Wulfe; Ted Xiao; Jonathan Heewon Yang; Arefeh Yavary; Tony Z. Zhao; Christopher Agia; Rohan Baijal; Mateo Guaman Castro; Daphne Chen; Qiuyu Chen; Trinity Chung; Jaimyn Drake; Ethan Paul Foster; Jensen Gao; David Antonio Herrera; Minho Heo; Kyle Hsu; Jiaheng Hu; Donovon Jackson; Charlotte Le; Yunshuang Li; Roy Lin; Zehan Ma; Abhiram Maddukuri; Suvir Mirchandani; Daniel Morton; Tony Nguyen; Abigail O'Neill; Rosario Scalise; Derick Seale; Victor Son; Stephen Tian; Emi Tran; Andrew E. Wang; Yilin Wu; Annie Xie; Jingyun Yang; Patrick Yin; Yunchu Zhang; Osbert Bastani; Glen Berseth; Jeannette Bohg; Ken Goldberg; Abhinav Gupta; Abhishek Gupta; Dinesh Jayaraman; Joseph J. Lim; Jitendra Malik; Roberto Mart\'ın-Mart\'ın; Subramanian Ramamoorthy; Dorsa Sadigh; Shuran Song; Jiajun Wu; Michael C. Yip; Yuke Zhu; Thomas Kollar; Sergey Levine; Chelsea Finn

DROID expands robot demonstration coverage by moving a standardized manipulation platform through many real workspaces. Its 76k successful episodes support diffusion-policy co-training that improves average executed-task success, while retaining task-specific failures. The central evidence concerns data diversity and supervised robot control, rather than learned world prediction (e-resource, e-policy, e-main).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

TD-MPC2: Scalable, Robust World Models for Continuous Control

Nicklas Hansen; Hao Su; Xiaolong Wang

TD-MPC2 learns a decoder-free latent world model and uses short-horizon planning, completed by a learned terminal value, to choose continuous actions. Its contribution is a coordinated robustness recipe and a task-conditioned architecture that supports larger multitask models; the evidence concerns simulated control, not generated-video quality or physical deployment. [e03, e07, e08, e09, e11]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning

Jianlan Luo; Zheyuan Hu; Charles Xu; You Liang Tan; Jacob Berg; Archit Sharma; Stefan Schaal; Chelsea Finn; Abhishek Gupta; Sergey Levine

SERL integrates demonstration-assisted off-policy RL, reward specification, learned resets and compliant robot control into a practical software stack. It learns physically executed manipulation policies from images and proprioception; the reported system reaches 100/100 successes on three tasks. The evidence supports an effective integrated implementation, without isolating which component causes each gain [e02, e03, e09, e12, e20].

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

VidMan: Exploiting Implicit Dynamics from Video Diffusion Model for Effective Robot Manipulation

Youpeng Wen; Junfan Lin; Yi Zhu; Jianhua Han; Hang Xu; Shen Zhao; Xiaodan Liang

VidMan adapts a video diffusion transformer into a language-conditioned robot policy. Video prediction on OXE supplies pretrained dynamics features; gated adapters and a diffusion action head turn those features into action sequences without generating future images during control. Evidence supports improved simulated manipulation and offline prediction, while the latent inverse-dynamics interpretation remains an architectural inductive bias.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Improved Distribution Matching Distillation for Fast Image Synthesis

Tianwei Yin; Michaël Gharbi; Taesung Park; Richard Zhang; Eli Shechtman; Frédo Durand; William T. Freeman

DMD2 turns pretrained image diffusion models into one- or few-step generators by stabilizing distribution matching, adding real-image adversarial supervision, and training multi-step models on their own intermediate samples. Critic accuracy and training-input choice matter alongside the loss itself. The one-step SDXL implementation retains a small regression warm-start. [e02, e03, e05, e18]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes

Rishabh Agarwal; Nino Vieillard; Yongchao Zhou; Piotr Stanczyk; Sabela Ramos Garea; Matthieu Geist; Olivier Bachem

Generalized Knowledge Distillation (GKD) improves a language-model student by querying its teacher on prefixes the student actually generates. It separates the choice of training sequences from the choice of token-distribution divergence. T5 experiments support using on-policy data, while showing task-dependent divergence preferences and extra sampling cost. This is a training method for text generation, with no demonstrated world-model control loop.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Revisiting Sparse Rewards for Goal-Reaching Reinforcement Learning

Gautham Vasan; Yan Wang; Fahim Shahriar; James Bergstra; Martin Jägersand; A. Rupam Mahmood

This empirical SAC study finds that constant −1 rewards with goal termination can produce better final reaching behavior despite slower learning. Initial goal encounters help select useful reset timeouts. Its physical-robot evidence concerns learned pixel-based control, with no learned world model or imagined rollout (formulations; simulation-final; timeout-diagnostic; architecture; conclusion).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

OccWorld: Learning a 3D Occupancy World Model for Autonomous Driving

Wenzhao Zheng; Weiliang Chen; Yuanhui Huang; Borui Zhang; Yueqi Duan; Jiwen Lu

OccWorld compresses semantic 3D occupancy into discrete scene tokens and jointly forecasts those tokens and an ego-motion token. Its shared spatial-temporal transformer supports both scene forecasting and trajectory planning. Oracle occupancy gives substantially stronger results than camera-derived occupancy, while long-horizon error and collision rates limit the driving conclusions.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots

Soroush Nasiriany; Abhiram Maddukuri; Lance Zhang; Adeet Parikh; Aaron Lo; Abhishek Joshi; Ajay Mandlekar; Yuke Zhu

RoboCasa packages diverse kitchen simulation, human-curated tasks and scalable demonstration generation for imitation learning. Its strongest controlled trend is higher average atomic-task success with larger generated datasets; composite-task and real-robot results remain modest. The contribution is infrastructure for learning and evaluation, with BC policies as baselines [e02, e06, e08, e09, e10, e11].

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving

Xiaofeng Wang; Zheng Zhu; Guan Huang; Xinze Chen; Jiagang Zhu; Jiwen Lu

DriveDreamer learns a driving-scene generator from structured real-world observations, then predicts the missing future structure from actions. Auto-DM renders that structure into video, while an action decoder combines its features with action history. The experiments support generation quality, synthetic-data usefulness, and open-loop trajectory prediction; they do not establish closed-loop driving safety.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

One-step Diffusion with Distribution Matching Distillation

Tianwei Yin; Michaël Gharbi; Richard Zhang; Eli Shechtman; Frédo Durand; William T. Freeman; Taesung Park

Distribution Matching Distillation (DMD) converts a pretrained diffusion denoiser into a one-pass image generator. A frozen teacher and a continually updated model of generated images supply a score-difference gradient; an offline teacher-pair regression loss anchors structure and diversity. The reported image-quality gains over earlier one-step methods retain a gap to the teacher and require substantial distillation training.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Genie: Generative Interactive Environments

Jake Bruce; Michael D Dennis; Ashley Edwards; Jack Parker-Holder; Yuge Shi; Edward Hughes; Matthew Lai; Aditi Mavalankar; Richie Steigerwald; Chris Apps; Yusuf Aytar; Sarah Maria Elisabeth Bechtle; Feryal Behbahani; Stephanie C.Y. Chan; Nicolas Heess; Lucy Gonzalez; Simon Osindero; Sherjil Ozair; Scott Reed; Jingwei Zhang; Konrad Zolna; Jeff Clune; Nando De Freitas; Satinder Singh; Tim Rocktäschel

Genie learns a small latent control interface from unlabeled video, then generates future frames from an image prompt and user-selected codes. A separate video tokenizer, latent action model and dynamics transformer make this possible. Its ablations distinguish visual fidelity from action sensitivity; its CoinRun experiment transfers inferred actions into a separately trained imitation policy. This is evidence for interactive neural simulation and latent-action pretraining, with limited evidence for persistent worlds or general-purpose control.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Diffusion for World Modeling: Visual Details Matter in Atari

Eloi Alonso; Adam Jelley; Vincent Micheli; Anssi Kanervisto; Amos Storkey; Tim Pearce; François Fleuret

DIAMOND learns an action-conditioned diffusion simulator in pixels, then trains a separate recurrent policy on imagined experience. EDM preconditioning makes short denoising schedules practical. Atari returns support its usefulness, while offline 3D demonstrations expose memory and causal limitations (e03, e04, e05, e09, e14, e21).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Craftax: A Lightning-Fast Benchmark for Open-Ended Reinforcement Learning

Michael Matthews; Michael Beukman; Benjamin Ellis; Mikayel Samvelyan; Matthew Thomas Jackson; Samuel Coward; Jakob Nicolaus Foerster

Craftax makes a procedurally generated survival game fast enough for billion-step reinforcement-learning studies. Its JAX implementation combines symbolic observations with accelerator-resident simulation and training. The harder environment exposes a gap between collecting easy rewards and exploring deeper floors: recurrent PPO helps, but the tested exploration and curriculum methods leave substantial challenges unsolved.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Video generation models as world simulators

OpenAI

Sora unifies images and videos as patches of compressed visual latents, then uses a text-conditioned diffusion transformer to generate them. This official overview explains the representation and describes qualitative capabilities rather than a reproducible benchmark study. Its central world-simulation claim concerns emergent visual consistency; acknowledged failures in physics and object-state changes limit that interpretation. [e02, e03, e04, e10, e12]

Resource reviewedText reportFull text availableRead the report Primary source

CoPeD-Advancing Multi-Robot Collaborative Perception: A Comprehensive Dataset in Real-World Environments

Yang Zhou; Long Quang; Carlos Nieto-Granda; Giuseppe Loianno

CoPeD records real indoor and outdoor air-ground robot teams with overlapping camera and LiDAR views, pose estimates, and optional automatic perception labels. Its contribution is a heterogeneous collection and annotation pipeline. A graph-network demonstration suggests qualitative depth recovery under camera corruption; it does not establish numerical benchmark gains or action-execution performance.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Octo: An Open-Source Generalist Robot Policy

Dibya Ghosh; Homer Rich Walke; Karl Pertsch; Kevin Black; Oier Mees; Sudeep Dasari; Joey Hejna; Tobias Kreiman; Charles Xu; Jianlan Luo; You Liang Tan; Lawrence Yunliang Chen; Quan Vuong; Ted Xiao; Pannag R. Sanketi; Dorsa Sadigh; Chelsea Finn; Sergey Levine

Octo learns a reusable visuomotor policy from 800k robot demonstrations. Modular input tokens feed a shared transformer, whose readout drives a diffusion action head. Its strongest evidence is adaptation across six physical manipulation setups with about 100 demonstrations each; flexibility still requires finetuning and does not imply universal zero-shot control. [e02, e03, e04, e10]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Vision-Language Foundation Models as Effective Robot Imitators

Xinghang Li; Minghuan Liu; Hanbo Zhang; Cunjun Yu; Jie Xu; Hongtao Wu; Chilam Cheang; Ya Jing; Weinan Zhang; Huaping Liu; Hang Li; Tao Kong

RoboFlamingo adapts OpenFlamingo into a language-conditioned manipulation policy: a vision-language backbone interprets each camera observation, while a recurrent head remembers history and predicts robot commands. CALVIN results support this decomposition, but robot-only fine-tuning sharply reduces general vision-language performance. The evidence concerns simulated action execution, with no learned future-world predictor or real-robot deployment.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models

Siddharth Karamcheti; Suraj Nair; Ashwin Balakrishna; Percy Liang; Thomas Kollar; Dorsa Sadigh

Prismatic VLMs studies which ingredients matter when a pretrained vision encoder feeds patch embeddings to a language model. Its main contribution is a controlled design investigation, supported by an evaluation suite and author-reported training infrastructure. Single-stage optimization and DINOv2–SigLIP feature fusion improve aggregate performance, while raw tables reveal meaningful task tradeoffs. The resulting PRISM family strengthens visual question answering and localization. These are image-to-text and coordinate-prediction results; the paper does not evaluate executed robot actions or future-state prediction.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Open X-Embodiment Collaboration; Abby O'Neill; Abdul Rehman; Abhinav Gupta; Abhiram Maddukuri; Abhishek Gupta; Abhishek Padalkar; Abraham Lee; Acorn Pooley; Agrim Gupta; Ajay Mandlekar; Ajinkya Jain; Albert Tung; Alex Bewley; Alex Herzog; Alex Irpan; Alexander Khazatsky; Anant Rai; Anchit Gupta; Andrew Wang; Andrey Kolobov; Anikait Singh; Animesh Garg; Aniruddha Kembhavi; Annie Xie; Anthony Brohan; Antonin Raffin; Archit Sharma; Arefeh Yavary; Arhan Jain; Ashwin Balakrishna; Ayzaan Wahid; Ben Burgess-Limerick; Beomjoon Kim; Bernhard Schölkopf; Blake Wulfe; Brian Ichter; Cewu Lu; Charles Xu; Charlotte Le; Chelsea Finn; Chen Wang; Chenfeng Xu; Cheng Chi; Chenguang Huang; Christine Chan; Christopher Agia; Chuer Pan; Chuyuan Fu; Coline Devin; Danfei Xu; Daniel Morton; Danny Driess; Daphne Chen; Deepak Pathak; Dhruv Shah; Dieter Büchler; Dinesh Jayaraman; Dmitry Kalashnikov; Dorsa Sadigh; Edward Johns; Ethan Foster; Fangchen Liu; Federico Ceola; Fei Xia; Feiyu Zhao; Felipe Vieira Frujeri; Freek Stulp; Gaoyue Zhou; Gaurav S. Sukhatme; Gautam Salhotra; Ge Yan; Gilbert Feng; Giulio Schiavi; Glen Berseth; Gregory Kahn; Guangwen Yang; Guanzhi Wang; Hao Su; Hao-Shu Fang; Haochen Shi; Henghui Bao; Heni Ben Amor; Henrik I Christensen; Hiroki Furuta; Homanga Bharadhwaj; Homer Walke; Hongjie Fang; Huy Ha; Igor Mordatch; Ilija Radosavovic; Isabel Leal; Jacky Liang; Jad Abou-Chakra; Jaehyung Kim; Jaimyn Drake; Jan Peters; Jan Schneider; Jasmine Hsu; Jay Vakil; Jeannette Bohg; Jeffrey Bingham; Jeffrey Wu; Jensen Gao; Jiaheng Hu; Jiajun Wu; Jialin Wu; Jiankai Sun; Jianlan Luo; Jiayuan Gu; Jie Tan; Jihoon Oh; Jimmy Wu; Jingpei Lu; Jingyun Yang; Jitendra Malik; João Silvério; Joey Hejna; Jonathan Booher; Jonathan Tompson; Jonathan Yang; Jordi Salvador; Joseph J. Lim; Junhyek Han; Kaiyuan Wang; Kanishka Rao; Karl Pertsch; Karol Hausman; Keegan Go; Keerthana Gopalakrishnan; Ken Goldberg; Kendra Byrne; Kenneth Oslund; Kento Kawaharazuka; Kevin Black; Kevin Lin; Kevin Zhang; Kiana Ehsani; Kiran Lekkala; Kirsty Ellis; Krishan Rana; Krishnan Srinivasan; Kuan Fang; Kunal Pratap Singh; Kuo-Hao Zeng; Kyle Hatch; Kyle Hsu; Laurent Itti; Lawrence Yunliang Chen; Lerrel Pinto; Li Fei-Fei; Liam Tan; Linxi "Jim" Fan; Lionel Ott; Lisa Lee; Luca Weihs; Magnum Chen; Marion Lepert; Marius Memmel; Masayoshi Tomizuka; Masha Itkina; Mateo Guaman Castro; Max Spero; Maximilian Du; Michael Ahn; Michael C. Yip; Mingtong Zhang; Mingyu Ding; Minho Heo; Mohan Kumar Srirama; Mohit Sharma; Moo Jin Kim; Muhammad Zubair Irshad; Naoaki Kanazawa; Nicklas Hansen; Nicolas Heess; Nikhil J Joshi; Niko Suenderhauf; Ning Liu; Norman Di Palo; Nur Muhammad Mahi Shafiullah; Oier Mees; Oliver Kroemer; Osbert Bastani; Pannag R Sanketi; Patrick "Tree" Miller; Patrick Yin; Paul Wohlhart; Peng Xu; Peter David Fagan; Peter Mitrano; Pierre Sermanet; Pieter Abbeel; Priya Sundaresan; Qiuyu Chen; Quan Vuong; Rafael Rafailov; Ran Tian; Ria Doshi; Roberto Martín-Martín; Rohan Baijal; Rosario Scalise; Rose Hendrix; Roy Lin; Runjia Qian; Ruohan Zhang; Russell Mendonca; Rutav Shah; Ryan Hoque; Ryan Julian; Samuel Bustamante; Sean Kirmani; Sergey Levine; Shan Lin; Sherry Moore; Shikhar Bahl; Shivin Dass; Shubham Sonawani; Shubham Tulsiani; Shuran Song; Sichun Xu; Siddhant Haldar; Siddharth Karamcheti; Simeon Adebola; Simon Guist; Soroush Nasiriany; Stefan Schaal; Stefan Welker; Stephen Tian; Subramanian Ramamoorthy; Sudeep Dasari; Suneel Belkhale; Sungjae Park; Suraj Nair; Suvir Mirchandani; Takayuki Osa; Tanmay Gupta; Tatsuya Harada; Tatsuya Matsushima; Ted Xiao; Thomas Kollar; Tianhe Yu; Tianli Ding; Todor Davchev; Tony Z. Zhao; Travis Armstrong; Trevor Darrell; Trinity Chung; Vidhi Jain; Vikash Kumar; Vincent Vanhoucke; Vitor Guizilini; Wei Zhan; Wenxuan Zhou; Wolfram Burgard; Xi Chen; Xiangyu Chen; Xiaolong Wang; Xinghao Zhu; Xinyang Geng; Xiyuan Liu; Xu Liangwei; Xuanlin Li; Yansong Pang; Yao Lu; Yecheng Jason Ma; Yejin Kim; Yevgen Chebotar; Yifan Zhou; Yifeng Zhu; Yilin Wu; Ying Xu; Yixuan Wang; Yonatan Bisk; Yongqiang Dou; Yoonyoung Cho; Youngwoon Lee; Yuchen Cui; Yue Cao; Yueh-Hua Wu; Yujin Tang; Yuke Zhu; Yunchu Zhang; Yunfan Jiang; Yunshuang Li; Yunzhu Li; Yusuke Iwasawa; Yutaka Matsuo; Zehan Ma; Zhuo Xu; Zichen Jeff Cui; Zichen Zhang; Zipeng Fu; Zipeng Lin

Open X-Embodiment pools heterogeneous robot experience into a common dataset format and tests whether existing action policies benefit from it. The repository spans 22 embodiments, while RT-X experiments use a nine-manipulator mixture. The strongest transfer result concerns skills available in another robot’s demonstrations: RT-2-X reaches 75.8% success on the Google Robot, versus RT-2’s 27.3%. Gains depend on evaluation setting; generalization to unseen robots is untested.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

ACT-Bench: Towards Action Controllable World Models for Autonomous Driving

Hidehisa Arai; Keishi Ishihara; Tsubasa Takahashi; Yu Yamaguchi

ACT-Bench tests whether generated driving videos follow supplied ego trajectories. Its learned evaluator estimates a maneuver label and a path, separating instruction agreement from geometric error. TERRA improves the reported aggregate scores over Vista, but unequal conditioning frequency and unresolved evaluation details limit causal interpretation of that comparison.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Andreas Blattmann; Tim Dockhorn; Sumith Kulal; Daniel Mendelevitch; Maciej Kilian; Dominik Lorenz; Yam Levi; Zion English; Vikram Voleti; Adam Letts; Varun Jampani; Robin Rombach

Stable Video Diffusion turns an image diffusion backbone into a reusable video generator through curated video pretraining and high-quality finetuning. Its strongest methodological lesson is that pretraining data quality continues to matter after finetuning. Text-to-video, image-to-video and multi-view experiments test generative representations; they do not establish an action policy or closed-loop world model.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

ADriver-I: A General World Model for Autonomous Driving

Fan Jia; Weixin Mao; Yingfei Liu; Yucheng Zhao; Yuqing Wen; Chi Zhang; Xiangyu Zhang; Tiancai Wang

ADriver-I couples a multimodal language model that predicts speed and steering with an action-conditioned video diffusion model. Generated frames can feed the next control prediction. Its strongest evidence concerns supervised control prediction and short-horizon video quality; the recurrent demonstration does not establish reliable autonomous driving. The inspected source is the original arXiv v1.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

GAIA-1: A Generative World Model for Autonomous Driving

Anthony Hu; Lloyd Russell; Hudson Yeo; Zak Murez; George Fedoseev; Alex Kendall; Jamie Shotton; Gianluca Corrado

GAIA-1 generates driving videos by predicting discrete image tokens and rendering them with a separate video diffusion decoder. Video context, language, and supplied speed/curvature sequences can shape the future. Semantic tokenization and temporal subsampling make autoregressive modeling tractable; diffusion restores visual detail and frame rate. The evidence comprises token-prediction scaling, sampling diagnostics, and selected qualitative scenarios. These support a controllable video world model, while calibrated physical dynamics, executed driving performance, and benefits to policy learning remain unestablished in this paper.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Jinze Bai; Shuai Bai; Shusheng Yang; Shijie Wang; Sinan Tan; Peng Wang; Junyang Lin; Chang Zhou; Jingren Zhou

Qwen-VL turns a language model into a multilingual image-understanding system through a visual encoder, a position-aware compression adapter, and three training stages. It produces captions, answers and textual bounding boxes. Strong generalist benchmark results coexist with specialist gaps and limited causal evidence for individual design choices.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

LLaMA: Open and Efficient Foundation Language Models

Hugo Touvron; Thibaut Lavril; Gautier Izacard; Xavier Martinet; Marie-Anne Lachaux; Timothée Lacroix; Baptiste Rozière; Naman Goyal; Eric Hambro; Faisal Azhar; Aurelien Rodriguez; Armand Joulin; Edouard Grave; Guillaume Lample

LLaMA trains causal language models longer to improve capability at a chosen inference size. Public-data mixtures and established transformer refinements yield strong language benchmarks, with uneven gains across tasks. The study concerns text prediction and prompting; it does not introduce a world-action model.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Learning Universal Policies via Text-Guided Video Generation

Yilun Du; Mengjiao Yang; Bo Dai; Hanjun Dai; Ofir Nachum; Joshua B. Tenenbaum; Dale Schuurmans; Pieter Abbeel

UniPi turns a language instruction and current image into a video plan, refines its timing, then uses a separately trained inverse model to produce robot controls. Simulated manipulation results support compositional and multitask transfer. Internet pretraining improves generated real-scene plans, but its reported success is a classifier judgment on imagined final frames, not measured physical execution. [e02, e04, e06, e10, e14, e16]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

3D Gaussian Splatting for Real-Time Radiance Field Rendering

Bernhard Kerbl; Georgios Kopanas; Thomas Leimkühler; George Drettakis

3D Gaussian Splatting fits an explicit radiance field to calibrated photographs of a static scene. Anisotropic Gaussians, adaptive splitting/cloning and a differentiable tile rasterizer jointly enable fast novel-view synthesis. Its main tradeoff is high rendering speed with competitive image quality at substantial scene-memory cost; it learns neither actions nor temporal dynamics (e02, e09, e13).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

PaLM-E: An Embodied Multimodal Language Model

Danny Driess; Fei Xia; Mehdi S. M. Sajjadi; Corey Lynch; Aakanksha Chowdhery; Brian Ichter; Ayzaan Wahid; Jonathan Tompson; Quan Vuong; Tianhe Yu; Wenlong Huang; Yevgen Chebotar; Pierre Sermanet; Daniel Duckworth; Sergey Levine; Vincent Vanhoucke; Karol Hausman; Marc Toussaint; Klaus Greff; Andy Zeng; Igor Mordatch; Pete Florence

PaLM-E turns sensor observations into vectors interleaved with language embeddings, then generates answers or high-level robot instructions. Its central empirical finding is that broad multimodal training can improve robot-task learning with little task-specific data. Separate controllers execute its instructions. Results support a reusable embodied language backbone, with important differences between simulation scores, perception diagnostics and qualitative physical demonstrations.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

RT-1: Robotics Transformer for Real-World Control at Scale

Anthony Brohan; Noah Brown; Justice Carbajal; Yevgen Chebotar; Joseph Dabis; Chelsea Finn; Keerthana Gopalakrishnan; Karol Hausman; Alexander Herzog; Jasmine Hsu; Julian Ibarz; Brian Ichter; Alex Irpan; Tomas Jackson; Sally Jesmonth; Nikhil J. Joshi; Ryan Julian; Dmitry Kalashnikov; Yuheng Kuang; Isabel Leal; Kuang-Huei Lee; Sergey Levine; Yao Lu; Utsav Malla; Deeksha Manjunath; Igor Mordatch; Ofir Nachum; Carolina Parada; Jodilyn Peralta; Emily Perez; Karl Pertsch; Jornell Quiambao; Kanishka Rao; Michael S. Ryoo; Grecia Salazar; Pannag R. Sanketi; Kevin Sayed; Jaspiar Singh; Sumedh Sontakke; Austin Stone; Clayton Tan; Huong T. Tran; Vincent Vanhoucke; Steve Vega; Quan Vuong; Fei Xia; Ted Xiao; Peng Xu; Sichun Xu; Tianhe Yu; Brianna Zitkovich

RT-1 learns a language-conditioned robot policy by compressing six camera observations into task-relevant tokens and predicting discretized actions. It joins a 35M-parameter architecture with a large, varied demonstration collection. Real-robot results support compositional instruction generalization and heterogeneous-data transfer within kitchen manipulation. Long-horizon results require SayCan, and evaluation-count inconsistencies limit exact protocol reconstruction.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

DexGraspNet: A Large-Scale Robotic Dexterous Grasp Dataset for General Objects Based on Simulation

Ruicheng Wang; Jialiang Zhang; Jiayi Chen; Yinzhen Xu; Puhao Li; Tengyu Liu; He Wang

DexGraspNet supplies 1.32 million simulated ShadowHand grasps on 5,355 objects. Geometry-based optimization generates candidate poses, and stringent simulation filtering selects the dataset. Training two existing grasp predictors on it improves reported quality and diversity, but the benchmark uses a weaker success criterion than dataset construction. Precision and functional grasps remain poorly covered.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

BridgeData V2: A Dataset for Robot Learning at Scale

Homer Walke; Kevin Black; Abraham Lee; Moo Jin Kim; Max Du; Chongyi Zheng; Tony Zhao; Philippe Hansen-Estruch; Quan Vuong; Andre He; Vivek Myers; Kuan Fang; Chelsea Finn; Sergey Levine

BridgeData V2 makes varied manipulation experience reusable through a common robot platform, goal images and language labels. Six offline policy baselines demonstrate useful but uneven generalization. Its strongest controlled lesson is that broader skill coverage can help an unseen pick-and-place task; cross-morphology transfer remains untested.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

DREAMWALKER: Mental Planning for Continuous Vision-Language Navigation

Hanqing Wang; Wei Liang; Luc Van Gool; Wenguan Wang

DREAMWALKER builds an episodic waypoint graph, synthesizes future panoramic observations, and searches those imagined states before executing navigation actions. It reports 49% test success on VLN-CE, with a short planning horizon limiting both computation and prediction drift. The evidence supports simulated navigation benefits; conflicting implementation descriptions leave exact reproduction unresolved.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Scalable Diffusion Models with Transformers

William Peebles; Saining Xie

DiT replaces the U-Net denoiser in latent image diffusion with a transformer over spatial latent patches. Conditioning design and the amount of computation per denoising pass both matter: adaLN-Zero works best among the tested blocks, and larger backbones or smaller patches improve ImageNet generation. Its strongest 256×256 result uses long training and classifier-free guidance; the experiments establish image-generation performance, not action-conditioned dynamics or control.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Xingchao Liu; Chengyue Gong; Qiang Liu

Rectified flow learns distribution transport by regressing straight-line endpoint displacements onto intermediate states. Reflow trains on the learned flow's endpoint pairs to make coarse numerical integration more accurate; final distillation further improves one-step generation. The CIFAR-10 results establish a speed–quality tradeoff, while translation and domain adaptation demonstrate broader uses of the transport construction. This report concerns the verified September 2022 v1 preprint.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

MimicGen: A Data Generation System for Scalable Robot Learning using Human Demonstrations

Ajay Mandlekar; Soroush Nasiriany; Bowen Wen; Iretiayo Akinola; Yashraj Narang; Linxi Fan; Yuke Zhu; Dieter Fox

MimicGen expands a small human demonstration set by transforming object-relative motion segments, executing them in new scenes, and retaining successful trajectories for behavioral cloning. Its strongest evidence concerns useful manipulation data across reset distributions, objects and robot arms. Generation yield, learned-policy success and real-world performance are distinct outcomes (E03–E04, E11, E13–E15).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection

Hao Zhang; Feng Li; Shilong Liu; Lei Zhang; Hang Su; Jun Zhu; Lionel M. Ni; Heung-Yeung Shum

DINO turns image features into object boxes and classes through dynamic anchor queries. Contrastive denoising, mixed query initialization and adjacent-layer box supervision improve DETR-style detection, with an accuracy–computation tradeoff between four and five feature scales. This is the July 2022 v4 detector paper; it studies neither future-state prediction nor action execution. [e01, e04, e05, e08, e10]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Brianna Zitkovich; Tianhe Yu; Sichun Xu; Peng Xu; Ted Xiao; Fei Xia; Jialin Wu; Paul Wohlhart; Stefan Welker; Ayzaan Wahid; Quan Vuong; Vincent Vanhoucke; Huong T. Tran; Radu Soricut; Anikait Singh; Jaspiar Singh; Pierre Sermanet; Pannag R. Sanketi; Grecia Salazar; Michael S. Ryoo; Krista Reymann; Kanishka Rao; Karl Pertsch; Igor Mordatch; Henryk Michalewski; Yao Lu; Sergey Levine; Lisa Lee; Tsang-Wei Edward Lee; Isabel Leal; Yuheng Kuang; Dmitry Kalashnikov; Ryan Julian; Nikhil J. Joshi; Alex Irpan; Brian Ichter; Jasmine Hsu; Alexander Herzog; Karol Hausman; Keerthana Gopalakrishnan; Chuyuan Fu; Pete Florence; Chelsea Finn; Kumar Avinava Dubey; Danny Driess; Tianli Ding; Krzysztof Marcin Choromanski; Xi Chen; Yevgen Chebotar; Justice Carbajal; Noah Brown; Anthony Brohan; Montserrat Gonzalez Arenas; Kehang Han

RT-2 makes a pretrained vision-language model a robot policy by encoding low-level actions as text tokens and co-fine-tuning on demonstrations and web tasks. Its clearest gain is transferring visual and semantic knowledge to unfamiliar objects and instructions. Physical execution improves, but unfamiliar dynamics and new motor skills remain major boundaries.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Scaling Data Generation in Vision-and-Language Navigation

Zun Wang; Jialu Li; Yicong Hong; Yi Wang; Qi Wu; Mohit Bansal; Stephen Gould; Hao Tan; Yu Qiao

ScaleVLN turns scanned indoor environments into synthetic instruction–trajectory supervision for existing navigation agents. Graph connectivity, image recovery and training mixtures matter alongside scale. Standard DUET with ScaleVLN reaches 77% R2R test-unseen success; the headline 80% additionally uses EnvEdit and CLIP ViT-H/14. These are simulator navigation results, including continuous simulation.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

DayDreamer: World Models for Physical Robot Learning

Philipp Wu; Alejandro Escontrela; Danijar Hafner; Pieter Abbeel; Ken Goldberg

DayDreamer applies Dreamer to online learning on four physical robots. A learned latent dynamics model supplies imagined experience for actor-critic training while the policy continues collecting real data. Locomotion and manipulation improve within hours, but navigation only matches its baseline. The contribution is practical deployment across robot interfaces with shared hyperparameters; it does not establish a new algorithm or a general robot policy.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Transformers are Sample-Efficient World Models

Vincent Micheli; Eloi Alonso; François Fleuret

IRIS learns an Atari policy inside a world model that compresses images into discrete tokens and predicts their evolution with an autoregressive Transformer. Real experience trains the simulator; imagined rewards train a separate recurrent actor-critic. Its Atari 100k mean and interquartile mean exceed the reported no-search baselines, while its median trails SPR. Reconstruction errors, rare events and substantial training compute delimit the result.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Sigmoid Loss for Language Image Pre-Training

Xiaohua Zhai; Basil Mustafa; Alexander Kolesnikov; Lucas Beyer

SigLIP learns aligned image and text representations with a pairwise sigmoid objective. Its practical contribution combines independent pair losses, memory-efficient distributed evaluation, and optimization choices. Controlled experiments favor sigmoid most at small batches; increasing batch size indefinitely does not reliably improve transfer. SigLiT reuses a frozen vision encoder, whereas SigLIP trains both towers. These are representation-learning methods for classification and retrieval, not action predictors.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning

Bo Liu; Yifeng Zhu; Chongkai Gao; Yihao Feng; Qiang Liu; Yuke Zhu; Peter Stone

LIBERO tests whether a robot can learn successive manipulation tasks while retaining earlier skills. Its controlled task suites distinguish spatial, object and goal changes; its experiments reveal a tradeoff between learning new tasks quickly and preserving previous performance. The reported outcomes are simulated policy execution, with strong dependence on architecture, learning algorithm and evaluation protocol.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

OpenScene: The Largest Up-to-Date 3D Occupancy Prediction Benchmark in Autonomous Driving

OpenScene Contributors

OpenScene is a compact redistribution of nuPlan with additional occupancy labels and occupancy-grid motion annotations. The supplied README describes the data resource and its challenge uses; it does not specify or evaluate a learned world-action model. This report covers that overview only. The reported scale and storage reduction describe dataset packaging, not measured prediction or driving performance.

Resource reviewedText reportFull text availableRead the report Primary source

Zenseact Open Dataset: A large-scale and diverse multimodal dataset for autonomous driving

Mina Alibeigi; William Ljungbergh; Adam Tonderski; Georg Hess; Adam Lilja; Carl Lindström; Daria Motorniuk; Junsheng Fu; Jenny Widahl; Christoffer Petersson

Zenseact Open Dataset (ZOD) combines geographically dispersed driving keyframes with shorter and longer continuous recordings. Its contribution is calibrated multimodal data and perception annotations, with baseline experiments exposing weaknesses at long distances and on rare classes. The reviewed release supports research on perception, temporal reasoning, and ego-motion, but its sparse temporal labels and incomplete distant annotations limit what those experiments establish about tracking or autonomous control.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Consistency Models

Yang Song; Prafulla Dhariwal; Mark Chen; Ilya Sutskever

Consistency Models maps a noisy image at any diffusion time directly to the low-noise endpoint of its probability-flow trajectory. This supports one-network-evaluation generation and optional refinement. Consistency distillation (CD) learns from a pretrained diffusion teacher; consistency training (CT) learns from paired perturbations of real data. CD has the better reported one-step FIDs of these two training routes, while CT removes diffusion-teacher dependence. The experiments concern images, with no executed actions or world-model control.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks

Oier Mees; Lukás Hermann; Erick Rosete-Beas; Wolfram Burgard

CALVIN tests whether a language-conditioned robot policy can compose manipulation skills across five successive instructions and transfer across related simulated scenes. It supplies play data, sparse language labels, multimodal interfaces and state-based success checks. Its MCIL baseline learns actions from demonstrations through latent plans; it does not demonstrate learned future-world prediction. Static-camera MCIL reaches 53.9% short-horizon success in environment D but only 0.08% five-instruction completion under a different, neutral-start evaluation protocol (e02, e07, e08, e10).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Temporal Difference Learning for Model Predictive Control

Nicklas A Hansen; Hao Su; Xiaolong Wang

TD-MPC couples short-horizon continuous-action planning with a reward-oriented latent dynamics model and a learned terminal Q function. Training shapes predicted latents through reward, value and consistency losses; control executes the first optimized action and replans from feedback. The experiments support strong simulated-control performance with a task-dependent computation tradeoff, not a general-purpose environment simulator or demonstrated physical deployment.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Rapid Exploration for Open-World Navigation with Latent Goal Models

Dhruv Shah; Benjamin Eysenbach; Nicholas Rhinehart; Sergey Levine

RECON learns how a visual goal relates to the robot's current view, predicts actions and temporal distance, and samples latent subgoals for exploration. A topological image memory supports frontier selection and later route reuse. Real-robot experiments favor this combination over random-action exploration, but missing appendices and limited statistical reporting constrain reproduction and causal interpretation.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

LoRA: Low-Rank Adaptation of Large Language Models

Edward J. Hu; Yelong Shen; Phillip Wallis; Zeyuan Allen-Zhu; Yuanzhi Li; Shean Wang; Lu Wang; Weizhu Chen

LoRA adapts a frozen language model by learning low-rank additive weight updates. These updates can be merged into the original dense layers for inference. This v2 draft demonstrates competitive language understanding and generation with far fewer trainable parameters, while its ablations show that update placement and task-specific rank matter.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Ego4D: Around the World in 3,000 Hours of Egocentric Video

Kristen Grauman; Andrew Westbury; Eugene Byrne; Zachary Chavis; Antonino Furnari; Rohit Girdhar; Jackson Hamburger; Hao Jiang; Miao Liu; Xingyu Liu; Miguel Martin; Tushar Nagarajan; Ilija Radosavovic; Santhosh Kumar Ramakrishnan; Fiona Ryan; Jayant Sharma; Michael Wray; Mengmeng Xu; Eric Zhongcong Xu; Chen Zhao; Siddhant Bansal; Dhruv Batra; Vincent Cartillier; Sean Crane; Tien Do; Morrie Doulaty; Akshay Erapalli; Christoph Feichtenhofer; Adriano Fragomeni; Qichen Fu; Abrham Gebreselasie; Cristina González; James Hillis; Xuhua Huang; Yifei Huang; Wenqi Jia; Weslie Khoo; Jáchym Kolár; Satwik Kottur; Anurag Kumar; Federico Landini; Chao Li; Yanghao Li; Zhenqiang Li; Karttikeya Mangalam; Raghava Modhugu; Jonathan Munro; Tullie Murrell; Takumi Nishiyasu; Will Price; Paola Ruiz Puentes; Merey Ramazanova; Leda Sari; Kiran K. Somasundaram; Audrey Southerland; Yusuke Sugano; Ruijie Tao; Minh Vo; Yuchen Wang; Xindi Wu; Takuma Yagi; Ziwei Zhao; Yunyi Zhu; Pablo Arbeláez; David Crandall; Dima Damen; Giovanni Maria Farinella; Christian Fuegen; Bernard Ghanem; Vamsi Krishna Ithapu; C. V. Jawahar; Hanbyul Joo; Kris Kitani; Haizhou Li; Richard A. Newcombe; Aude Oliva; Hyun Soo Park; James M. Rehg; Yoichi Sato; Jianbo Shi; Mike Zheng Shou; Antonio Torralba; Lorenzo Torresani; Mingfei Yan; Jitendra Malik

Ego4D-3K turns long recordings of human daily activity into a resource for retrieving past events, interpreting present interactions and anticipating future behavior. Its contribution is collection and task design: 3,670 hours from 931 wearers, with uneven additional modalities. The supplied paper establishes dataset scope and evaluation targets, but delegates model scores and implementation details to absent appendices.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Contrastive Learning as Goal-Conditioned Reinforcement Learning

Benjamin Eysenbach; Tianjun Zhang; Sergey Levine; Russ R. Salakhutdinov

Contrastive RL turns discrimination between actual future states and random states into a goal-reaching critic. State-action and goal encoders produce an inner-product score; a separate actor learns actions that increase it. The exponentiated optimal score is proportional to a goal-averaged policy value. Simulated online and offline results support this approach, with restrictive proof assumptions and unavailable appendices limiting verification.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Socially Compliant Navigation Dataset (SCAND): A Large-Scale Dataset of Demonstrations for Social Navigation

Haresh Karnan; Anirudh Nair; Xuesu Xiao; Garrett Warnell; Sören Pirk; Alexander Toshev; Justin W. Hart; Joydeep Biswas; Peter Stone

SCAND records human joystick control alongside robot sensing so social navigation can be learned from demonstrated behavior. It contains 138 trajectories, 8.7 hours and approximately 40 km (reported also as 25 miles), collected by four demonstrators using Jackal and Spot. Its preliminary Spot experiments distinguish demonstrators, imitate future paths and velocity commands, and obtain favorable participant ratings against move_base. These establish useful demonstration data and limited policy validation, rather than comprehensive social-navigation performance (e02–e14).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

High-Resolution Image Synthesis With Latent Diffusion Models

Robin Rombach; Andreas Blattmann; Dominik Lorenz; Patrick Esser; Björn Ommer

Latent diffusion separates perceptual compression from generative modeling: a pretrained autoencoder supplies a spatial latent grid, a diffusion UNet learns to denoise it, and a decoder returns an image. Cross-attention adds text, class or layout conditioning. Experiments favor moderate compression and demonstrate image synthesis and editing, while exposing a reconstruction bottleneck and tradeoffs between distributional quality, coverage and pixel fidelity.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

TartanDrive: A Large-Scale Dataset for Learning Off-Road Dynamics Models

Samuel Triest; Matthew Sivaprakasam; Sean J. Wang; Wenshan Wang; Aaron M. Johnson; Sebastian A. Scherer

TartanDrive supplies human-driven ATV interactions for learning how terrain, vehicle motion and commands jointly determine future pose. Its benchmark fuses visual and proprioceptive observations into an action-conditioned latent dynamics model. Additional modalities improve reported prediction error, particularly on uneven terrain; whether those improvements enable autonomous navigation remains untested.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

BEVDet: High-performance Multi-camera 3D Object Detection in Bird-Eye-View

Junjie Huang; Guan Huang; Zheng Zhu; Yun Ye; Dalong Du

BEVDet converts calibrated multi-camera images into a shared bird's-eye-view representation and detects present 3D objects there. Its main contribution is making this familiar modular pipeline train effectively: image-space augmentation and BEV-space augmentation address different parts of the network, while Scale-NMS suppresses small-object duplicates. The experiments establish a useful nuScenes accuracy–speed tradeoff under the reported protocols, with weaker attribute prediction and important timing qualifications. This is scene perception, without an action-conditioned dynamics model or demonstrated driving controller.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

BERT: A Review of Applications in Natural Language Processing and Understanding

M. V. Koroteev

Koroteev surveys how pretrained BERT representations support classification, extractive summarization, text evaluation, adversarial testing, multilingual learning and scientific language models. A shared encoder is reused through different data, objectives and downstream interfaces. This is a synthesis of earlier studies, without a new unified model or experiment. Its readable evidence includes SciBERT task comparisons and a training ablation whose interpretation requires separating loss changes from input construction.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Mastering Atari with Discrete World Models

Danijar Hafner; Timothy Lillicrap; Mohammad Norouzi; Jimmy Ba

DreamerV2 learns a categorical latent dynamics model from images, then trains a separate actor and critic entirely on imagined latent trajectories. Its reported Atari advantage concerns aggregate performance under a 200M-frame, sticky-action protocol. Discrete latents, KL balancing, and image-derived representation gradients are important ingredients; the experiments do not establish universal game mastery or physical robot control.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Pathdreamer: A World Model for Indoor Navigation

Jing Yu Koh; Honglak Lee; Yinfei Yang; Jason Baldridge; Peter Anderson

Pathdreamer imagines what an indoor agent might see along a supplied future trajectory. It completes semantic segmentation and depth first, then renders RGB, using accumulated 3D point clouds to preserve visual context. Its predictions improve a separate instruction-following navigation system on unseen Matterport3D buildings. The central tradeoff is useful visual foresight without reliable recovery of the actual unseen layout; navigation still assumes a ground-truth graph of feasible movements.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Deep Reinforcement Learning at the Edge of the Statistical Precipice

Rishabh Agarwal; Max Schwarzer; Pablo Samuel Castro; Aaron Courville; Marc Bellemare

This paper turns multi-task RL evaluation into an uncertainty-aware statistical workflow: keep individual runs, bootstrap within tasks, inspect score distributions and summarize them with robust metrics. Its experiments show how sampling noise and evaluation protocols can change apparent progress. The contribution is evaluation methodology, with no new agent architecture or controller. [e03, e04, e06, e08, e11]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Learning Transferable Visual Models From Natural Language Supervision

Alec Radford; Jong Wook Kim; Chris Hallacy; Aditya Ramesh; Gabriel Goh; Sandhini Agarwal; Girish Sastry; Amanda Askell; Pamela Mishkin; Jack Clark; Gretchen Krueger; Ilya Sutskever

CLIP learns image and text embeddings by identifying paired examples in Internet data, then turns natural-language class descriptions into a visual classifier. Its strongest model reaches 76.2% ImageNet accuracy without fitting on ImageNet training labels. Transfer breadth and natural-shift robustness are substantial, but specialized tasks, evaluation co-adaptation and missing supplementary protocols limit the conclusions.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations

Tongzhou Mu; Zhan Ling; Fanbo Xiang; Derek Yang; Xuanlin Li; Stone Tao; Zhiao Huang; Zhiwei Jia; Hao Su

ManiSkill benchmarks simulated manipulation of unseen objects within familiar categories. Its contribution is a curated set of articulated environments, successful demonstrations and controlled evaluation tracks. Point-cloud policies can learn one fixed object reasonably well, but the reported baselines struggle across object variation. This is a benchmark and demonstration pipeline, with no learned future-world predictor proposed.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

NuPlan: A closed-loop ML-based planning benchmark for autonomous vehicles

Holger Caesar; Juraj Kabzan; Kok Seang Tan; Whye Kit Fong; Eric Wolff; Alex Lang; Luke Fletcher; Oscar Beijbom; Sammy Omari

nuPlan proposes a benchmark that turns recorded driving into repeated planner–controller–simulator evaluation. Its central distinction is between predicting a trajectory that resembles a recording and executing a plan whose consequences change the next planning state. This version describes a planned 1500-hour dataset, open-loop and two closed-loop tasks, and candidate planning metrics; it reports no evaluated nuPlan planner or ablation. The contribution is an evaluation design, with agent realism and metric aggregation still unresolved (e-protocol, e-tasks, e-scale, e-metrics, e-proposal).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Beyond the Nav-Graph: Vision-and-Language Navigation in Continuous Environments

Jacob Krantz; Erik Wijmans; Arjun Majumdar; Dhruv Batra; Stefan Lee

VLN-CE turns language-guided graph traversal into low-level navigation through reconstructed indoor spaces. Its recurrent policies choose movements from egocentric RGBD observations without supplied location or topology. The strongest recipe attains 32% success and 0.30 SPL on unseen continuous environments; this is simulated navigation performance, with graph-based comparisons subject to conversion caveats.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Model-Based Reinforcement Learning for Atari

Lukasz Kaiser; Mohammad Babaeizadeh; Piotr Milos; Blazej Osinski; Roy H. Campbell; Konrad Czechowski; Dumitru Erhan; Chelsea Finn; Piotr Kozakowski; Sergey Levine; Afroz Mohiuddin; Ryan Sepassi; George Tucker; Henryk Michalewski

SimPLe learns an action-conditioned video-and-reward simulator from limited Atari experience, trains a separate PPO policy inside it, and repeatedly collects new real-game data. Discrete stochastic latents and short simulated rollouts make imperfect prediction useful for control. Its contribution is sample-efficient policy learning, with substantial computation and instability; the supplied revision also acknowledges competitive later model-free baselines.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Alexander Ku; Peter Anderson; Roma Patel; Eugene Ie; Jason Baldridge

Room-Across-Room (RxR) pairs multilingual navigation instructions with routes and synchronized human camera poses. Its sampling and annotation procedures make route adherence central to evaluation. The recurrent navigation baselines benefit from complementary demonstrations but remain far below human followers: test-standard success averages 25.9% for monolingual agents versus 93.9% for humans. [e-dataset, e-sampling, e-annotation, e-test]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

nuScenes: A Multimodal Dataset for Autonomous Driving

Holger Caesar; Varun Bankiti; Alex H. Lang; Sourabh Vora; Venice Erin Liong; Qiang Xu; Anush Krishnan; Yu Pan; Giancarlo Baldan; Oscar Beijbom

nuScenes pairs surround-view camera, lidar and radar recordings with 3D object annotations and maps, then defines detection and tracking evaluations. Its central lesson is that data scale, temporal input and the matching rule all affect perceived progress. The baselines estimate objects and tracks; they do not demonstrate a learned driving controller. [e01, e03, e08, e10, e12, e13, e14]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

RLBench: The Robot Learning Benchmark & Learning Environment

Stephen James; Zicong Ma; David Rovick Arrojo; Andrew J. Davison

RLBench standardizes simulated manipulation around 100 hand-designed tasks, multimodal observations and waypoint-generated demonstrations. Its central evaluation distinction is between new configurations of familiar tasks and genuinely held-out tasks. This v1 paper establishes an environment and few-shot protocol, with descriptive task statistics but no learned-policy performance comparison (e02, e05, e11, e13, e16).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Scalability in Perception for Autonomous Driving: Waymo Open Dataset

Pei Sun; Henrik Kretzschmar; Xerxes Dotiwalla; Aurelien Chouard; Vijaysai Patnaik; Paul Tsui; James Guo; Yin Zhou; Yuning Chai; Benjamin Caine; Vijay Vasudevan; Wei Han; Jiquan Ngiam; Hang Zhao; Aleksei Timofeev; Scott Ettinger; Maxim Krivokon; Amy Gao; Aditya Joshi; Yu Zhang; Jonathon Shlens; Zhifeng Chen; Dragomir Anguelov

Waymo Open Dataset supplies synchronized camera/LiDAR sequences, independent spatial annotations and perception benchmarks. Its PointPillars results expose long-range difficulty, geographic transfer loss and gains from more training sequences. The evidence concerns offline detection and tracking; printed metric and protocol inconsistencies constrain exact reproduction.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning

Fisher Yu; Haofeng Chen; Xin Wang; Wenqi Xian; Yingying Chen; Fangchen Liu; Vashisht Madhavan; Trevor Darrell

BDD100K organizes driving videos into annotation sets with different costs and output structures. Its experiments ask whether abundant simpler labels improve scarce, harder perception tasks. Detection supervision helps instance segmentation, tracking, and semantic segmentation, but transfer depends on the auxiliary task and available data. The combined MOTS result improves aggregate accuracy while retaining substantial errors; it establishes perception transfer, not autonomous driving competence. [task-overview, instseg-transfer, mot-transfer, semantic-transfer, mots-transfer]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

Colin Raffel; Noam Shazeer; Adam Roberts; Katherine Lee; Sharan Narang; Michael Matena; Yanqi Zhou; Wei Li; Peter J. Liu

T5 turns classification, question answering, summarization, and translation into conditional text generation. A controlled transfer-learning study selects an encoder–decoder Transformer and economical denoising, then combines these choices with C4, multi-task pre-training, task-specific fine-tuning, and scale. The evidence supports a reusable language backbone; its strongest benchmark results use a substantially different regime from the baseline ablations.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Leveraging Procedural Generation to Benchmark Reinforcement Learning

Karl Cobbe; Christopher Hesse; Jacob Hilton; John Schulman

Procgen benchmarks reinforcement learning with 16 procedurally generated games. Its central intervention is to vary the levels encountered by an agent and evaluate unseen levels separately. PPO baselines expose substantial overfitting; wider convolutional models improve held-out returns, while fixed level sequences can produce impressive training scores with little transfer (e02, e07–e10).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning

Tianhe Yu; Deirdre Quillen; Zhanpeng He; Ryan Julian; Karol Hausman; Chelsea Finn; Sergey Levine

Meta-World tests whether reinforcement learning can share manipulation skills across 50 simulated task families and adapt to held-out families. Shared robot and control interfaces make transfer plausible, while task diversity exposes failures hidden by goal-only benchmarks. Multi-headed SAC helps on ten known tasks, but broad multi-task learning and few-shot adaptation remain difficult. These are simulated policy-execution results, with important protocol and reporting ambiguities.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Language Models are Few-Shot Learners

Tom Brown; Benjamin Mann; Nick Ryder; Melanie Subbiah; Jared D Kaplan; Prafulla Dhariwal; Arvind Neelakantan; Pranav Shyam; Girish Sastry; Amanda Askell; Sandhini Agarwal; Ariel Herbert-Voss; Gretchen Krueger; Tom Henighan; Rewon Child; Aditya Ramesh; Daniel Ziegler; Jeffrey Wu; Clemens Winter; Chris Hesse; Mark Chen; Eric Sigler; Mateusz Litwin; Scott Gray; Benjamin Chess; Jack Clark; Christopher Berner; Sam McCandlish; Alec Radford; Ilya Sutskever; Dario Amodei

GPT-3 tests whether scaling autoregressive text pretraining makes a fixed model more useful when tasks are specified by examples in its context. The 175-billion-parameter model achieves strong completion and factual-QA results without downstream gradient updates, but gains vary sharply by task. Prompt format, evaluation split, pretraining exposure and substantial compute are central to interpreting the evidence. [e-settings, e-architecture, e-lambada, e-qa, e-superglue, e-contamination, e-compute]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Dream to Control: Learning Behaviors by Latent Imagination

Danijar Hafner; Timothy Lillicrap; Jimmy Ba; Mohammad Norouzi

Dreamer learns a visual world model from experience, then trains a separate actor and value model through imagined latent trajectories. Bootstrapped values extend the learning signal beyond a short rollout; environment interaction uses the learned actor with freshly inferred states. The continuous-control results support this approach, subject to unequal baseline budgets and incomplete baseline task coverage. [e-flow, e-returns, e-updates, e-scores]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Behaviour Suite for Reinforcement Learning

Ian Osband; Yotam Doron; Matteo Hessel; John Aslanides; Eren Sezener; Andre Saraiva; Katrina McKinney; Tor Lattimore; Csaba Szepesvári; Satinder Singh; Benjamin Van Roy; Richard S. Sutton; David Silver; Hado van Hasselt

Behaviour Suite (bsuite) tests reinforcement-learning agents through small, scalable experiments whose environments, interaction budgets and analyses are fixed together. Its memory and deep-exploration examples expose different weaknesses in recurrent actor critic, DQN and Bootstrapped DQN. This reading concerns the ICLR 2020/arXiv v3 paper and its bsuite2019 experiments, including unresolved main-text/appendix protocol differences (identity; protocol; memory-result; deepsea-result).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Learning to summarize with human feedback

Nisan Stiennon; Long Ouyang; Jeffrey Wu; Daniel Ziegler; Ryan Lowe; Chelsea Voss; Alec Radford; Dario Amodei; Paul F. Christiano

A pretrained summarizer first learns from reference summaries, then a separate reward model learns which candidate summaries humans prefer. PPO uses that learned score, constrained toward the supervised policy, to improve generation. A 1.3B feedback policy obtains 61% preference against reference summaries, versus 43% for a roughly ten-times-larger supervised policy. News transfer is promising, but excessive reward optimization eventually reduces human preference. The contribution concerns learning a text-generation objective, rather than modeling physical dynamics or executing robot actions (e-pipeline, e-ppo, e-tldr, e-transfer, e-overoptimization).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

RoboNet: Large-Scale Multi-Robot Learning

Sudeep Dasari; Frederik Ebert; Stephen Tian; Suraj Nair; Bernadette Bucher; Karl Schmeckpeper; Siddharth Singh; Sergey Levine; Chelsea Finn

RoboNet pools robot experience to make visual control transferable. Pretraining improves adaptation with a few hundred target-robot trajectories, but relevant subsets can outperform the broader pool. Its central contribution is a shared dataset evaluated through two distinct control algorithms.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark

Xiaohua Zhai; Joan Puigcerver; Alexander Kolesnikov; Pierre Ruyssen; Carlos Riquelme; Mario Lucic; Josip Djolonga; Andre Susano Pinto; Maxim Neumann; Alexey Dosovitskiy; Lucas Beyer; Olivier Bachem; Michael Tschannen; Marcin Michalski; Olivier Bousquet; Sylvain Gelly; Neil Houlsby

VTAB evaluates how well representations adapt to diverse unseen classification tasks with limited labels. Its ImageNet study finds that combining supervision with self-supervision improves transfer, while tuning budget and the choice between fine-tuning and frozen features materially affect conclusions. The supplied revision also contains unresolved protocol and numerical inconsistencies, documented below.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

General Evaluation for Instruction Conditioned Navigation using Dynamic Time Warping

Gabriel Ilharco; Vihan Jain; Alexander Ku; Eugene Ie; Jason Baldridge

nDTW scores how closely an agent follows an ordered reference trajectory using distance-aware dynamic time warping. SDTW additionally requires goal success. Human path rankings and R2R/R4R reward comparisons support their usefulness, while implementation ambiguities and limited experimental reporting constrain reproduction (E02–E10, E14).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

When to Trust Your Model: Model-Based Policy Optimization

Michael Janner; Justin Fu; Marvin Zhang; Sergey Levine

Model-based policy optimization (MBPO) uses a learned dynamics-and-reward ensemble to generate extra training transitions for soft actor-critic. Each short rollout begins at a real replay-buffer state, limiting accumulated prediction error while permitting many policy updates per environment interaction. The central question is how much model usage improves learning before model bias dominates. The paper combines conditional return bounds, empirical model-generalization diagnostics and MuJoCo experiments. Its strongest explicit comparison matches SAC's Ant performance using 300,000 rather than 3 million environment steps; it does not establish an equivalent reduction in computation time.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

LIC-Fusion: LiDAR-Inertial-Camera Odometry

Xingxing Zuo; Patrick Geneva; Woosik Lee; Yong Liu; Guoquan Huang

LIC-Fusion estimates motion by combining inertial propagation, sparse camera tracks and LiDAR edge/plane constraints in one MSCKF estimator, while refining sensor transforms and clock offsets online. Its strongest outdoor evidence is a reported average ATE of 4.06 m, versus 10.75 m for MSCKF and 23.08 m for LOAM. Indoor results favor fusion on the harder sequences but favor LOAM on A and B; endpoint-only evaluation limits the conclusion.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Habitat: A Platform for Embodied AI Research

Manolis Savva; Jitendra Malik; Devi Parikh; Dhruv Batra; Abhishek Kadian; Oleksandr Maksymets; Yili Zhao; Erik Wijmans; Bhavana Jain; Julian Straub; Jia Liu; Vladlen Koltun

Habitat separates 3D scene ingestion, fast sensor simulation and task evaluation. Its PointGoal experiments show that longer PPO training changes the comparison with one SLAM baseline, with depth-only agents strongest under idealized sensing. Cross-dataset results expose training-distribution effects. The contribution is experimental infrastructure, not a learned world-action predictor (e02, e03, e04, e13, e14, e16, e18).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Argoverse: 3D Tracking and Forecasting With Rich Maps

Ming-Fang Chang; John Lambert; Patsorn Sangkloy; Jagjeet Singh; Slawomir Bak; Andrew Hartnett; De Wang; Peter Carr; Simon Lucey; Deva Ramanan; James Hays

Argoverse makes detailed maps usable alongside autonomous-driving observations: lane geometry guides trajectory forecasts, while ground height and driveable area constrain tracking. Its contribution is a dataset and benchmark with illustrative baselines. Reported improvements support map-aware perception, but unequal forecast hypothesis budgets and orientation-insensitive tracking metrics limit causal conclusions (e-resource, e-maps, e-tracking-protocol, e-multimodal, e-forecast-results).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Diversity is all you need: Learning skills without a reward function

Benjamin Eysenbach; Abhishek Gupta; Julian Ibarz; Sergey Levine

DIAYN learns a repertoire of stochastic, state-feedback skills without an external task reward. A discriminator rewards states that reveal which skill generated them, while action entropy and a fixed skill prior discourage narrow behavior and neglected skills. The repertoire supports reward-based adaptation, hierarchical control and state-only imitation, but task usefulness depends on what states become distinguishable (e-objective, e-loop, e-transfer, e-hierarchy, e-imitation).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

Antoine Miech; Dimitri Zhukov; Jean-Baptiste Alayrac; Makarand Tapaswi; Ivan Laptev; Josef Sivic

HowTo100M uses subtitle timing to turn narrated web videos into weakly paired clips and text. A shallow joint embedding learns correspondence at scale, supporting retrieval and action-step localization. Its strongest transfer is instructional; movie retrieval exposes the limits of off-the-shelf transfer.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Learning Latent Dynamics for Planning from Pixels

Danijar Hafner; Timothy Lillicrap; Ian Fischer; Ruben Villegas; David Ha; Honglak Lee; James Davidson

PlaNet learns an action-conditioned latent dynamics model from images and rewards, then searches that model online to choose continuous actions. Its recurrent state-space model combines deterministic memory with stochastic latent states. The experiments establish sample-efficient simulated control, with uneven gains across tasks. Latent overshooting is a separate proposed training objective; the final RSSM agent omits it.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Towards Accurate Generative Models of Video: A New Metric & Challenges

Thomas Unterthiner; Sjoerd van Steenkiste; Karol Kurach; Raphael Marinier; Marcin Michalski; Sylvain Gelly

Fréchet Video Distance (FVD) compares distributions of real and generated videos using pretrained action-recognition features and Gaussian statistics. The paper validates sensitivity to temporal corruption and agreement with human preferences, then introduces StarCraft 2 Videos (SCV) to expose failures in motion, interaction and memory. Its central tradeoff is a useful distribution-level score whose interpretation depends on the embedding, sample count and evaluation protocol.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

On Evaluation of Embodied Navigation Agents

Peter Anderson; Angel Chang; Devendra Singh Chaplot; Alexey Dosovitskiy; Saurabh Gupta; Vladlen Koltun; Jana Kosecka; Jitendra Malik; Roozbeh Mottaghi; Manolis Savva; Amir R. Zamir

This working-group report defines how to compare embodied navigation agents: specify goals, sensors and prior exploration, evaluate an explicit completion action using traversable distance, and summarize successful navigation with SPL. Its contribution is an evaluation contract and standard scenarios across four environment datasets. It proposes no trained agent and reports no measured agent comparison; the numerical SPL examples explain the metric.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

World Models

David Ha; Jürgen Schmidhuber

World Models separates visual compression, recurrent prediction and reward-driven control. A VAE and an MDN-RNN learn from random gameplay; a small controller then uses their representations. CarRacing demonstrates useful predictive features during simulator-based policy training. VizDoom demonstrates policy training inside a learned latent simulator followed by transfer to the original game. Its temperature sweep exposes the central risk: a controller can exploit model errors and obtain excellent imagined returns while failing after transfer.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

DeepMind Control Suite

Yuval Tassa; Yotam Doron; Alistair Muldal; Tom Erez; Yazhe Li; Diego de Las Casas; David Budden; Abbas Abdolmaleki; Josh Merel; Andrew Lefrancq; Timothy Lillicrap; Martin Riedmiller

DeepMind Control Suite standardizes simulated continuous-control tasks, interfaces and bounded rewards so reinforcement-learning agents can be compared across physical domains. Its baseline study favors distributed D4PG in aggregate, while revealing that pixel observations, training budgets and encoder updates strongly affect the conclusions.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Vision-and-Language Navigation: Interpreting Visually-Grounded Navigation Instructions in Real Environments

Peter Anderson; Qi Wu; Damien Teney; Jake Bruce; Mark Johnson; Niko Sünderhauf; Ian D. Reid; Stephen Gould; Anton van den Hengel

R2R turns natural-language route following into a reproducible benchmark over photographs of real buildings. Its simulator executes viewpoint transitions; an attention-based recurrent policy learns to choose navigation actions. Student-forcing improves baseline success, but the large seen/unseen gap exposes limited environmental generalization (e02, e13–e16).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

The Unreasonable Effectiveness of Deep Features as a Perceptual Metric

Richard Zhang; Phillip Isola; Alexei A. Efros; Eli Shechtman; Oliver Wang

LPIPS measures reference-to-image-patch dissimilarity through normalized deep features, optionally calibrated with human preferences. BAPPS tests agreement with people across synthetic distortions and real algorithm outputs. Pretrained features outperform common low-level metrics; modest calibration transfers more reliably than full fine-tuning. The evidence concerns local visual similarity with a reference patch, rather than general image quality or action competence.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Sim-to-Real Transfer of Robotic Control with Dynamics Randomization

Xue Bin Peng; Marcin Andrychowicz; Wojciech Zaremba; Pieter Abbeel

Dynamics randomization trains a recurrent pushing policy across varied simulated physics, then deploys it directly on a Fetch arm. Its memory can use interaction history, while a training-only critic receives privileged simulator parameters. Reported real success is 0.89 ± 0.06 over 28 trials. This is evidence for physical policy transfer under the tested setup, with different simulation and real evaluation horizons; it is not a learned future-prediction model.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

The Kinetics Human Action Video Dataset

Will Kay; Joao Carreira; Karen Simonyan; Brian Zhang; Chloe Hillier; Sudheendra Vijayanarasimhan; Fabio Viola; Tim Green; Trevor Back; Paul Natsev; Mustafa Suleyman; Andrew Zisserman

Kinetics turns short, visually verified YouTube actions into a large classification benchmark. Its contribution is the collection and cleanup pipeline plus baseline diagnostics: appearance and motion contribute differently across classes, and combined streams lead the reported Kinetics comparison. This is the original 400-class paper, with non-exhaustive labels and unresolved count inconsistencies (e01, e02, e03, e04, e08, e14, e18, e19).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Matterport3D: Learning from RGB-D Data in Indoor Environments

Angel X. Chang; Angela Dai; Thomas A. Funkhouser; Maciej Halber; Matthias Nießner; Manolis Savva; Shuran Song; Andy Zeng; Yinda Zhang

Matterport3D converts panoramic RGB-D scans of 90 buildings into registered images, reconstructed surfaces and semantic annotations. Its five separate perception baselines test whether broad viewpoint coverage and cleaner geometry improve matching, retrieval, normal prediction and semantic understanding. The results support useful supervision and transfer, while leaving action execution and learned dynamics untested (e-acquisition, e-annotation, e-keypoint, e-overlap, e-normal-results, e-region, e-voxel).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Factor Graphs for Robot Perception

Frank Dellaert; Michael Kaess

This selected-chapter reading follows Dellaert and Kaess from a probabilistic SLAM model to square root smoothing and mapping. Given noisy observations, known data associations and suitable priors, the estimator jointly adjusts robot poses and landmarks. Conditioning turns a generative Bayesian network into a factor graph; Gaussian measurement errors then lead to nonlinear least squares, repeated linearization and matrix factorization. The central lesson is how local sensor constraints become a global estimation problem. The reviewed chapters provide derivations and diagrams, not an empirical evaluation of learned world models or robot control.

Partially reviewedIllustrated editionPartial text availableRead the report Primary source

Neural Discrete Representation Learning

Aaron van den Oord; Oriol Vinyals; Koray Kavukcuoglu

VQ-VAE learns a discrete codebook by quantizing encoder outputs, training reconstruction with a straight-through gradient, and fitting an autoregressive prior afterward. It supports compact image and speech representations and action-conditioned latent video generation. Its evidence combines likelihood bounds, a phoneme probe and qualitative examples; it does not evaluate action selection or executed control.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

The "Something Something" Video Database for Learning and Evaluating Visual Common Sense

Raghav Goyal; Samira Ebrahimi Kahou; Vincent Michalski; Joanna Materzynska; Susanne Westphal; Heuna Kim; Valentin Haenel; Ingo Fruend; Peter Yianilos; Moritz Mueller-Freitag; Florian Hoppe; Christian Thurau; Ingo Bax; Roland Memisevic

Something-Something turns everyday human–object interactions into supervised video-understanding tasks: workers enact action templates, select objects and fill the templates' noun slots. The 2017 paper describes 108,499 clips and 174 template classes, with real-versus-pretended actions intended to discourage superficial recognition. Its experiments predict template labels; they do not demonstrate future-video generation, physical-property estimation or robot control. Standard baselines remain weak, and the strongest combined encoder is evaluated only on simplified label subsets.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium

Martin Heusel; Hubert Ramsauer; Thomas Unterthiner; Bernhard Nessler; Sepp Hochreiter

This paper couples two contributions: separate generator/discriminator learning rates for GAN training, and FID for comparing generated images with real data through Inception features. It presents conditional local-equilibrium convergence arguments and selected improvements for DCGAN and WGAN-GP. The experiments support useful training and evaluation procedures; they do not certify that practical fixed-rate runs satisfy every theorem assumption.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation

Charles R. Qi; Hao Su; Kaichun Mo; Leonidas J. Guibas

PointNet learns directly from unordered 3D points: shared pointwise networks extract features, coordinatewise max pooling summarizes a shape, and classification or segmentation heads predict labels. Learned alignment improves accuracy, while a critical-point analysis explains conditional robustness. The evidence establishes an efficient geometric recognition representation, with qualifications about input modalities, missing supplements and corruption settings.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

CARLA: An Open Urban Driving Simulator

Alexey Dosovitskiy; Germán Ros; Felipe Codevilla; Antonio M. López; Vladlen Koltun

CARLA supplies controllable urban simulation and evaluation signals for driving policies. Its original benchmark compares modular control, conditional imitation learning and A3C, exposing large failures on a new town even when unseen weather is manageable. Completion and infraction metrics reveal different strengths.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Deep Reinforcement Learning from Human Preferences

Paul F. Christiano; Jan Leike; Tom Brown; Miljan Martic; Shane Legg; Dario Amodei

The paper trains a separate reward predictor from human comparisons of short behavior clips, then uses that predictor to train a reinforcement-learning policy. Repeated feedback connects reward learning to the behavior the policy actually visits. MuJoCo and Atari experiments demonstrate useful control with sparse human supervision, with task-dependent failures and limited replication of human-feedback runs. The mechanism learns preferences and actions, without predicting future world states.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

CIDEr: Consensus-Based Image Description Evaluation

Ramakrishna Vedantam; C. Lawrence Zitnick; Devi Parikh

CIDEr evaluates how closely an image caption matches the consensus of human reference descriptions. It combines TF-IDF weighted n-gram similarities with a triplet-based human validation protocol and densely captioned evaluation datasets. Its gains depend on reference coverage and pair type; the paper separately introduces CIDEr-D to discourage gaming.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Advances and applications of occupancy models

Larissa L Bailey; Darryl I MacKenzie; James D Nichols

This abstract describes a discussion of ecological occupancy estimation: inferring species occurrence and occupancy dynamics while allowing for missed detections or species misidentification. Its central message is that flexible models increase the investigator’s responsibility to define sites, sampling occasions, a period of static occurrence and detection criteria. Those definitions affect biological interpretation. The authors use an amphibian–pathogen system to illustrate the issues. Only the abstract and bibliographic record were supplied; this report cannot assess the detailed models or the applications’ empirical findings.

Partially reviewedText reportPartial text availableRead the report Primary source

Auto-Encoding Variational Bayes

Diederik P Kingma; Max Welling

Auto-Encoding Variational Bayes learns a probabilistic image generator together with a recognition network that approximates its otherwise intractable latent posterior. A differentiable transformation of independent noise makes the variational objective trainable by stochastic gradients. The paper separates the general SGVB estimator from its AEVB learning algorithm and neural variational-autoencoder example. MNIST and Frey Face experiments support improved lower-bound optimization over wake-sleep, while estimated marginal likelihood reveals a more qualified small-data comparison with Monte Carlo EM. These are static-image density-modeling results, not action or dynamics experiments. [e-problem, e-sgvb, e-vae, e-fig2, e-fig3, e-future]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

The Arcade Learning Environment: An Evaluation Platform for General Agents

Marc G. Bellemare; Yavar Naddaf; Joel Veness; Michael Bowling

The Arcade Learning Environment makes Atari games a common testbed for general-purpose agents. Its contribution combines an emulator interface, a separation between design games and evaluation games, and metrics for comparing incompatible score scales. Linear reinforcement-learning agents and planners using the exact simulator establish useful baselines, while exposing representation failures, sparse rewards, and misleading aggregate rankings.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

A benchmark for the evaluation of RGB-D SLAM systems

Jrgen Sturm; Nikolas Engelhard; Felix Endres; Wolfram Burgard; Daniel Cremers

This benchmark pairs Kinect color and registered depth images with externally measured camera poses, then evaluates estimated trajectories through local relative pose error (RPE) and globally aligned absolute trajectory error (ATE). Its contribution is calibrated data and a comparison protocol. The strongest quantitative evidence concerns sensor and ground-truth calibration; algorithm comparisons are illustrative plots. Robot-mounted recordings were manually driven and do not demonstrate autonomous control. [e02, e05, e07, e09, e11, e12, e13, e14, e15]

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Are we ready for autonomous driving? The KITTI vision benchmark suite

Andreas Geiger; Philip Lenz; Raquel Urtasun

KITTI turns synchronized driving recordings into separately scored stereo, optical-flow, odometry and object-perception benchmarks. Its contribution is calibrated reference data, selection rules and diagnostic metrics. The 2012 results expose failures under large motion and difficult surfaces; they do not measure autonomous driving success (e02, e05, e12, e13).

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Monte-Carlo Planning in Large POMDPs

David Silver; Joel Veness

POMCP plans under partial observability by sharing simulator trajectories between a history-based UCT search and an unweighted particle belief. It selects an action, observes the environment, and reuses the compatible subtree. The supplied simulator is not learned. Experiments show scalability in Rocksample, Battleship and PocMan, with domain-dependent gains from search and preferred actions; the formal guarantee assumes the true belief state.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments

Satanjeev Banerjee; Alon Lavie

METEOR evaluates English machine translations through explicit word alignment to human references, recall-weighted matching, and a fragmentation penalty. This 2005 paper tests agreement with human translation judgments and diagnoses the contributions of scoring and lexical matching. Its strong system-level correlation and much lower sentence-level correlations answer different evaluation questions.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Probabilistic Robotics

Sebastian Thrun; Wolfram Burgard; Dieter Fox

This introductory chapter explains why a robot should maintain a distribution over possible states and consider how its actions change uncertainty. Global localization shows sensing narrowing competing location hypotheses while motion spreads them; coastal navigation shows why a longer, informative route can be preferable. The chapter also identifies computational and approximation costs. It provides conceptual foundations and qualitative illustrations, with detailed algorithms deferred to later chapters.

Partially reviewedIllustrated editionPartial text availableRead the report Primary source

Image Quality Assessment: From Error Visibility to Structural Similarity

Zhou Wang; Alan C. Bovik; Hamid R. Sheikh; Eero P. Simoncelli

SSIM evaluates an aligned image against a reference by comparing local luminance, contrast and normalized structure, then averaging the local scores into MSSIM. It requires no learned predictor. On this paper's JPEG/JPEG2000 database, MSSIM tracks human quality ratings better than PSNR, Sarnoff and UQI under the reported criteria. The evidence supports perceptual assessment of compressed still images; it does not validate a general measure of semantic correctness, temporal dynamics or control performance.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

ROUGE: A Package for Automatic Evaluation of Summaries

Chin-Yew Lin

ROUGE measures reference-summary overlap through n-grams, longest common subsequences, weighted subsequences and ordered word pairs. Lin validates these deterministic scores against human content-coverage judgments on DUC summarization tasks. Strong system-level correlations in single-document summarization weaken or change across headline and multi-document settings, making metric variant, preprocessing and reference policy essential parts of the evaluation.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Real-time humanoid motion generation through ZMP manipulation based on inverted pendulum control

Tomomichi Sugihara; Yoshihiko Nakamura; Hirochika Inoue

This classical humanoid controller turns a desired center-of-gravity velocity into a feasible zero moment point, then into whole-body joint references. An inverted-pendulum approximation supplies the dynamics, while a constrained COG Jacobian distributes motion across joints. Simulated stepping survives an unspecified impact, but the paper provides no hardware validation, measured runtime or comparative error statistics.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Bleu: a Method for Automatic Evaluation of Machine Translation

Kishore Papineni; Salim Roukos; Todd Ward; Wei-Jing Zhu

BLEU evaluates translated text through clipped reference n-gram overlap and a corpus brevity penalty. Its Chinese-to-English study reports strong agreement with aggregate human judgments across three machine systems and two human translators. The evidence supports efficient corpus-level comparison under a shared reference protocol, with explicit limits on sentence scoring and cross-reference comparisons.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Planning and acting in partially observable stochastic domains

Leslie Pack Kaelbling; Michael L. Littman; Anthony R. Cassandra

This foundational planning paper converts uncertainty about a hidden state into a probability distribution, solves for reward-maximizing behavior over that belief space, and sometimes compiles the solution into a finite-memory controller. Its witness algorithm constructs exact finite-horizon value functions without enumerating every observation-contingent tree. The tiger examples explain information gathering and controller memory; scalability and dependence on a correct supplied model remain central limits.

Full-text reviewedIllustrated editionFull text availableRead the report Primary source

Reinforcement Learning: An Introduction

Richard S. Sutton; Andrew G. Barto

This selected reading of Sutton and Barto's online first edition explains how decisions can improve through value estimates learned from interaction and through planning with a model. Chapters 1 and 3 establish rewards, policies, state representations, and MDPs; Chapter 4 develops dynamic programming; Chapter 9 joins these ideas in Dyna and examines where planning computation should go. The central connection to world modeling is operational: predicted transitions become simulated experience that improves a policy through value updates. This is a foundational textbook treatment, with illustrative computational studies, rather than a single proposed neural architecture. Missing equation and figure images limit the review to verifiable prose and surviving experimental descriptions.

Partially reviewedText reportPartial text availableRead the report Primary source

Model predictive control: Theory and practice—A survey

Carlos E. García; David M. Prett; Manfred Morari

This abstract presents a survey of model predictive control as a family of controllers that directly use an explicit, separately identifiable model. Its organizing questions concern relationships among controller designs, constraint handling, performance objectives, nonlinear applicability and robustness. The authors emphasize easier adjustment for robustness while denying any inherent robustness advantage over classical feedback. These are abstract-level descriptions and author claims; the supplied material does not expose the arguments, examples or measurements needed to evaluate them (ev-definition, ev-survey-scope, ev-objectives, ev-nonlinear, ev-robustness, ev-source-scope).

Partially reviewedText reportPartial text availableRead the report Primary source

Advancing AI for the physical world

Microsoft Research

Microsoft Research's official overview introduces Rho-alpha, a Phi-derived robotics model that combines vision-language understanding with tactile sensing to produce control signals for bimanual manipulation. Its described training mixes physical and simulated trajectories with visual question answering data. The page offers demonstration descriptions and a development agenda, without quantitative evaluation or an implementable architecture. Human-assisted recovery is described; learning from that feedback remains a goal (e-interface, e-training, e-feedback, e-dual-arm, e-scope).

Resource reviewedText reportFull text availableRead the report Primary source
Browse all 558 illustrated reports

Original source visuals accompany explanations of each work and its world–action connection. Every reading records its evidence and source limitations.

Current RGB and depth frames share a frozen VAE encoder; a single diffusion transformer denoises action and future RGB-D slots.
READING 01 · 5 ORIGINAL VISUALS

SA-WAM

SA-WAM fits metric depth into a frozen video tokenizer and jointly predicts geometry and actions, improving tested manipulation while adding a geometric input and prediction burden.

Read illustrated report →
Three phases show frozen-model candidate selection, execution revealing prediction error, and online learning that affects later choices.
READING 02 · 6 ORIGINAL VISUALS

WCD

WCD spends extra test-time compute choosing among frozen-model rollouts, then learns from delayed prediction errors to improve future choices.

Read illustrated report →
Video, latent-action and action branches share attention during training; mask matrices show future-video slots absent at inference.
READING 03 · 6 ORIGINAL VISUALS

LAWA

LAWA replaces test-time future-video generation with continuous latent transition prediction, preserving tested control performance at intermediate inference cost.

Read illustrated report →
A blue 30-layer Slow video transformer sends video key/value features from its first 12 layers to a green 12-layer Fast action transformer; each branch has separate observation and state inputs.
READING 04 · 5 ORIGINAL VISUALS

ZimaBlue

ZimaBlue grounds video priors in robot actions, then reuses a slow world model’s cached features for fast corrective control, trading fresher feedback against stale guidance and distillation loss.

Read illustrated report →
A shared video–action transformer and a tactile expert appear beside a policy-to-simulator-to-evaluator chain whose values feed either policy optimization or maximum-value action selection.
READING 05 · 6 ORIGINAL VISUALS

Motus2

Motus2 makes action generation, future simulation, and progress evaluation conditional interfaces of one model, enabling optional planning and policy improvement while memory and tactile feedback add distinct deployment costs.

Read illustrated report →
A condition stream encodes language and current images, while a noise stream encodes noisy actions and future-image latents; their Q/K/V outputs meet in masked attention before action and future-state prediction.
READING 06 · 5 ORIGINAL VISUALS

DELE-w0.5

DELE-w0.5 trains actions alongside compact future-state latents but removes future-observation tokens at deployment, gaining a shorter inference sequence while leaving the causal benefit of future-state supervision unablated.

Read illustrated report →
Two-stage diagram: frozen DINOv3 features train an inverse-dynamics latent and world decoder; a backbone then predicts a latent subgoal that conditions a Flow-DiT action expert.
READING 07 · 5 ORIGINAL VISUALS

AcrossWAM 1.0

AcrossWAM separates transition prediction, future-feature decoding and action generation so a compact backbone can be substituted, but the reported substitution still adapts the decoder and action stack.

Read illustrated report →
Current RGB, pointmap and DINO representations enter a five-billion-parameter visual transformer with a separate one-billion-parameter action expert; two masks govern observation presence and future generation.
READING 08 · 5 ORIGINAL VISUALS

FLEX-π

FLEX-π trains one action model to read optional RGB, semantic and geometric futures, letting inference spend more computation on imagined context while preserving an action-only route.

Read illustrated report →
Four architecture panels compare ViNT, NWM, VertiFormer and Hydra; Hydra fuses action, pose and image history and routes predicted intent into three colored discrete codebooks.
READING 09 · 5 ORIGINAL VISUALS

Hydra

Hydra searches learned discrete visual, pose and action intents before continuous execution, trading expensive candidate rendering for latent costs whose uncertainty signals can miss semantic hazards.

Read illustrated report →
SG-WAM architecture: a VLM forecasts two semantic feature branches, trained against future-frame teachers, and conditions joint video and action experts through cross-attention.
READING 10 · 6 ORIGINAL VISUALS

SG-WAM

Predict what the instruction means for the future scene, then let that foresight guide video and action together.

Read illustrated report →
Four representations progress from a text prompt through 3D bounding boxes and a coarse 3D proxy to a complete 3D scene; a rightward arrow labels increasing control, acquisition difficulty and train-test mismatch.
READING 11 · 5 ORIGINAL VISUALS

Programmable World Model

An executable box world maintains persistent facts, while a video renderer completes their appearance; strong global state scores leave entity-specific fidelity and component contributions unresolved.

Read illustrated report →
Two frozen DINOv3 encoders feed cross-attention blocks with crossed key/value connections, then separate action/state-conditioned predictors and a dashed dual-view training-loss path.
READING 12 · 6 ORIGINAL VISUALS

DUET-DINO

Cross-attention lets side- and wrist-view predictors inform one another before latent goal matching, improving orientation-sensitive planning at a substantial CEM inference cost.

Read illustrated report →
Three-row FolDeX overview: conceptual sim-to-real and rigid-to-deformable gaps; recovery, task, scene, and embodiment tracks; dataset, evaluation pipeline, rollout repository, and leaderboard connected by arrows.
READING 13 · 5 ORIGINAL VISUALS

FolDeX

FolDeX turns complete physical garment folding into a controlled test of robot-data reuse, with promising recovery-augmentation results but incomplete scoring and transfer specifications.

Read illustrated report →
Pipeline with solid training arrows from video generation through HOI reconstruction to RL grounding, and dashed inference arrows through the tracking controller to real-world execution.
READING 14 · 6 ORIGINAL VISUALS

GALATEA

GALATEA preserves hand and object motion from generated video plans, then learns physical execution through simulation tracking, trading a structured, reusable controller for dependence on reconstruction quality and reference-following recovery.

Read illustrated report →
HAM takes predicted frames, current and anchor images, and action chunks through visual and action encoders, four fusion blocks, and a sigmoid hallucination head.
READING 15 · 6 ORIGINAL VISUALS

HaWMPO

A learned hallucination score improves some imagined-rollout policy updates by discounting unreliable rewards, but the benefit depends on penalty strength and the score remains a proxy for physical validity.

Read illustrated report →
Current and future observations pass through shared encoder weights; current features feed joint action–future prediction, while the stopped future latent supplies a training target.
READING 16 · 6 ORIGINAL VISUALS

JEPA Policy

A shared two-pass policy uses demonstrated future latents to improve action learning, with little measured inference overhead but additional training cost and task-dependent diagnostic reliability.

Read illustrated report →
Flow diagram linking simulator recordings to a compact world model, imagined actor–critic learning and executed control, with an episode-calibration margin branch and a separate public tactile-recording study.
READING 17 · 6 ORIGINAL VISUALS

Compact Visuotactile Lifting

Aligned touch improves force forecasts, but reward alignment improves lifting without overcoming the separate limits of dynamics, calibration and force regulation.

Read illustrated report →
Two-panel pipeline: a single image branches into action-conditioned videos, SLAM reconstructs candidate geometry, and a policy commits a selected branch. The lower example highlights the straight trajectory among left, straight and right alternatives.
READING 18 · 4 ORIGINAL VISUALS

Valerant

Valerant turns a frozen action-conditioned video model into a 3D map-building search procedure, with geometric checks compensating for generated views that can drift or pass through walls.

Read illustrated report →
Calibration and history frames feed video latents, their poses join future-action embeddings, and both streams condition a DiT that outputs future video.
READING 19 · 6 ORIGINAL VISUALS

SyncWorld

Paired visual calibration teaches a video world model how commands map to motion in a new setup, enabling useful short-horizon simulation while leaving action selection dependent on a separate policy and judge.

Read illustrated report →
A shared WorldAgen block receives color-coded instruction, image, robot-state, action and placeholder tokens. Two right-hand boxes depict action prediction and future-observation prediction.
READING 20 · 6 ORIGINAL VISUALS

WorldAgen

WorldAgen adapts a shared action-and-observation predictor through exploratory observation supervision, gaining simulated task success at the cost of target-environment interaction and additional training.

Read illustrated report →
Four panels show separate video and action generators sharing attention, a five-source pretraining mixture, a joint timestep plane, and an 80-dimensional action layout.
READING 21 · 6 ORIGINAL VISUALS

OpenWAM

OpenWAM-α couples a video prior and dedicated action generator through joint denoising, improving embodied transfer while retaining pronounced weaknesses under visual disturbance.

Read illustrated report →
Visual corruption examples and five temporal operators above a left-to-right pipeline from degraded observations through encoding, prediction, CEM preference and environment execution.
READING 22 · 5 ORIGINAL VISUALS

Beyond Task Success

Paired diagnostics reveal where sensing disturbances attenuate or persist in world-model planning, but internal sensitivity does not by itself establish a change in executed task success.

Read illustrated report →
Three panels show multimodal RSSM pretraining, a reusable predictive prior, and a policy receiving observations, commanded axis, and WSM features before updating joint targets.
READING 23 · 6 ORIGINAL VISUALS

WM-Craftnet

A predictive visuotactile memory can improve dexterous control through policy conditioning, while its dependence on pretraining and limited severe-disturbance recovery constrain the generalization claim.

Read illustrated report →
Two grouped bar charts show failure-stage distributions for action-only versus joint WAM policies and completion drops after perturbing near-, mid- and long-horizon future latents.
READING 24 · 6 ORIGINAL VISUALS

ProWAM

ProWAM uses recurrent execution progress to select useful imagined-future tokens for action generation, improving manipulation while relying on demonstration-derived progress and contact priors.

Read illustrated report →
Two-part WHIRL diagram: observations feed a frozen behavior prior and residual policy; their actions are composed for robot execution with human intervention, and labeled replay trains the world model, critic and actor.
READING 25 · 6 ORIGINAL VISUALS

WHIRL

WHIRL turns human takeovers into actor-side risk predictions, improving dexterous learning in the reported runs while inheriting the labeling operator's habits.

Read illustrated report →
Implementation table listing intervention sampling, hard-negative mining, task outcomes, simulators, encoder and auxiliary heads, sensitivity ranges, and compute.
READING 26 · 5 ORIGINAL VISUALS

CLWM

CLWM improves simulated planning by supervising the outcome geometry of imagined action sequences, at the cost of privileged training labels and additional branching.

Read illustrated report →
Three camera views enter CoAE; the head image and instruction enter a VLM. A DiT predicts dense and sparse visual states in one denoising pass, then a separate IDM iteratively produces an action chunk.
READING 27 · 6 ORIGINAL VISUALS

GE-Act 2.0

A one-step visual future makes separate video and action pretraining connectable, but useful co-training requires selecting futures compatible with the demonstrated action.

Read illustrated report →
Three expert blocks show future visual/tactile prediction, planned actions, and a tactile expert whose delta corrects the plan through a tactile–action key/value cache.
READING 28 · 6 ORIGINAL VISUALS

TacPAC

TacPAC turns a cached prediction of expected contact into a reference for fast tactile action correction, while keeping the expensive prediction fixed until the next chunk.

Read illustrated report →
Three-stage diagram: image-based scene lifting and declarative configuration feed a physical transition, fixed geometry, observation generator and gated appearance loop, followed by output video frames.
READING 29 · 6 ORIGINAL VISUALS

TourPhysics

TourPhysics keeps simulation authoritative while diffusion fills appearance, gaining persistent visual control at the cost of a declared, incomplete scene hypothesis and non-real-time generation.

Read illustrated report →
GIFT overview with a VLA above and a video/action WAM below, connected to geometry distillation, affordance and goal decoders; a full token-color and frozen/trainable legend appears at right.
READING 30 · 6 ORIGINAL VISUALS

GIFT

Training visual features to retain geometry, interactions and goals improves matched robot policies without auxiliary inference inputs, at the cost of structured training labels and with uneven robustness gains.

Read illustrated report →
Six integration types with checkmarks for representation learning, VLA policies and world modeling, followed by coupling principles, realizations and intended benefits.
READING 31 · 6 ORIGINAL VISUALS

Toward Unified Robot Learning

The survey argues that reliable robot learning depends on useful feedback among representations, policies and predictive models, while its diagnostic shows how plausible video can lose hidden objects and physical constraints.

Read illustrated report →
WISE pipeline: real images and an instruction feed a VLA; a scheduler skips or invokes world-model rollouts; imagined videos receive rewards, reliability filtering, group advantages, and a policy update.
READING 32 · 6 ORIGINAL VISUALS

WISE

WISE spends world-model computation on selected interaction states to improve a VLA action head, trading exhaustive imagination for dependence on learned state selection and reliable trajectory scoring.

Read illustrated report →
Four-by-four attention permission matrix with query rows, key columns, crossed forbidden cells, and a complete legend for ego-state, reference-video, future-action and future-video tokens.
READING 33 · 6 ORIGINAL VISUALS

SV-WAM

SV-WAM uses future-video supervision to train a surround-view planner whose causal mask removes future-video computation at deployment, while a differentiable footprint penalty improves benchmark road compliance.

Read illustrated report →
Two panels connect a slow VL-JEPA motion predictor to FiLM in an Emu3-based fast autoregressive model. The slow panel shows future optical flows and latents; the fast panel shows instruction, image and action inputs and next-image/action outputs.
READING 34 · 6 ORIGINAL VISUALS

Drive-HWM

A cached forecast of motion latents guides a per-step driving policy, with reported NAVSIM gains offset by unresolved result discrepancies and incomplete implementation details.

Read illustrated report →
Three observed robot images feed encoder E. Blue arrows train action-conditioned latent prediction, red arrows recover the executed action from consecutive encoded states, and green arrows recover physical state from the preceding and current encodings.
READING 35 · 6 ORIGINAL VISUALS

Physically Grounded JEPA

Training a JEPA world model to recover physical state as well as actions improves its reported image-goal planning, while its latent trajectories become less straight on average.

Read illustrated report →
Six-column production diagram: assets, validation, distributed scheduling, PIE trajectory generation, MRQ replay/rendering, and output management, above a reliability layer.
READING 36 · 6 ORIGINAL VISUALS

Unreal Engine Data Pipeline

Recording physics-resolved trajectories before offline rendering enables aligned multi-view data production, while the learning value of its perceptual curation remains unmeasured.

Read illustrated report →
LEAP architecture: a frozen image encoder feeds an action-conditioned autoregressive rollout; terminal latent and decoded-state costs combine into energy whose gradient returns to candidate actions.
READING 37 · 6 ORIGINAL VISUALS

LEAP

LEAP improves action planning through a frozen latent world model by requiring decoded terminal-state agreement, at the cost of numerical goal supervision and additional online optimization.

Read illustrated report →
Pre-training diagram freezes the vision encoder and shallow language layers, mixes caption and robot inputs, and trains OFT, GR00T and PI heads; fine-tuning unfreezes the backbone and initializes one head.
READING 38 · 6 ORIGINAL VISUALS

VLAct

VLAct preserves early vision-language features and trains several action decoders against shared targets, producing a backbone that transfers to fresh downstream heads without deploying the pre-training ensemble.

Read illustrated report →
Two transition diagrams contrast a generic conditional predictor with LEON's observable map, context function, operator update, forcing branch and learned readout.
READING 39 · 6 ORIGINAL VISUALS

LEON

LEON makes latent transitions explicit through shared evolution operators and additive forcing, improving VLA-JEPA while preserving near-baseline performance in LaWAM's inference-time prediction pathway.

Read illustrated report →
A multi-view camera encoder, structured text prompt, and embodiment ID feed a shared Action/Video DiT, with a shared visual-latent velocity head and an embodiment-specific action velocity head.
READING 40 · 6 ORIGINAL VISUALS

Riemann-1.0

Action-first autoregression lets one transformer serve as a robot policy and a visual simulator, while the reported evaluations leave the contribution of its ordering and staged supervision unresolved.

Read illustrated report →
Overview diagram showing task-diverse robot data and paired human–robot data feeding a video model and an action model, with text or human-video instructions below and evaluation examples at right.
READING 41 · 6 ORIGINAL VISUALS

Zero-WAM

Zero-WAM learns to translate a human video prompt into robot futures and inverse-dynamics actions, gaining unseen-task transfer through synthetic paired data and auxiliary future prediction while retaining substantial execution failures.

Read illustrated report →
RGB, tactile, action, state and language inputs feed a Wan DiT with future video, tactile and action outputs. An attention grid blocks video queries from tactile keys and boosts action queries toward tactile keys.
READING 42 · 6 ORIGINAL VISUALS

Tactile-WAM

Directional tactile attention improves contact-rich action generation and preserves an RGB-only visual trajectory, while the separate contributions of contact supervision and attention bias remain unresolved.

Read illustrated report →
Past RGB feeds four perception blocks, an object-associated Gaussian state, a policy and world model, then future splats and rendered RGB.
READING 43 · 5 ORIGINAL VISUALS

4DGS-WAM

4DGS-WAM forecasts actor motion and transports existing object Gaussians, gaining persistent scene structure while inheriting perception, camera and unseen-content limitations.

Read illustrated report →
SimWAM diagram with a video VAE and video DiT above an action encoder and action DiT, text and ego encoders below, isolated training and inference attention matrices in the center, and future video outputs marked removed during inference and RL.
READING 44 · 6 ORIGINAL VISUALS

SimWAM

SimWAM uses future-video prediction to train a direct trajectory policy, trading explicit test-time imagination for dependence on the quality of learned current-observation features.

Read illustrated report →
Two block diagrams: a world model takes observation and action and predicts an observation or state; a world action model takes observation and outputs action, with a parenthesized future-observation output.
READING 45 · 4 ORIGINAL VISUALS

From World Models to World Action Models

The tutorial separates prediction targets from action-generation interfaces, revealing why future modeling can support a robot policy without guaranteeing useful control.

Read illustrated report →
Observation, language, timestep, video and action tokens enter a frozen video–action Transformer; a trainable scheduler adds stopping and relative noise-progression heads above it.
READING 46 · 6 ORIGINAL VISUALS

SANTS

SANTS spends video-denoising computation according to the current state, trading small potential action-quality gains against the latency of further refinement.

Read illustrated report →
Training and inference diagrams with Video and Action DiTs, visual history, task/state inputs, a per-step router and gameplay/GUI action decoders.
READING 47 · 6 ORIGINAL VISUALS

GameWAM

GameWAM turns video co-training into native game control through shared visual context, short action commitments and mode-specific decoding, while sampled sources remain a control vulnerability.

Read illustrated report →
Four-part GaussianWAM diagram: frozen visual teachers feed a Gaussian field and target cache; auxiliary heads supervise current tokens in dual-expert and unified Transformer policies.
READING 48 · 6 ORIGINAL VISUALS

GaussianWAM

GaussianWAM uses an offline 3D Gaussian teacher to improve current-observation WAM representations, gaining robustness while preserving the deployed backbone and adding unmeasured teacher-preparation work.

Read illustrated report →
A current multiview observation branches into future RGB frames above and future point-cloud geometry with ego poses below.
READING 49 · 6 ORIGINAL VISUALS

GeoWAM

GeoWAM turns forecast geometry into a conditioning signal for a deterministic driving policy, with strong reported benchmark scores but no controlled test isolating the source of its gains.

Read illustrated report →
Training diagram with separate causal VAE passes, a masked shared video transformer and action flow head; a lower timeline reuses lookahead latents across action chunks.
READING 50 · 6 ORIGINAL VISUALS

GlanceWAM

GlanceWAM retains generated future guidance through a shared latent backbone and a fast action path, but its sparse-refresh deployment claim is limited by inconsistent scheduling descriptions.

Read illustrated report →
Green Video DiT blocks receive VAE video and text inputs; blue Action DiT blocks receive noisy actions, command and ego status. Vertical connections carry video features; auxiliary losses and a scored trajectory pool supply training supervision.
READING 51 · 6 ORIGINAL VISUALS

ReWorld

ReWorld explicitly trains a video-to-action representation pathway, improving reported video and simulated planning scores while requiring careful separation of curriculum effects from video self-guidance.

Read illustrated report →
The LDM learns semantic and motion targets on the left; video, clean latent queries and action experts share attention on the right; three training stages appear below.
READING 52 · 5 ORIGINAL VISUALS

LD4WAM

LD4WAM learns motion-grounded transition targets from human and robot videos, then predicts those targets between future-video generation and action, while retaining a dependence on the quality of imagined scenes.

Read illustrated report →
The student loop links environment observations, history, video, stop-gradient action generation and execution. Frozen Teacher Video z_T connects by an outer orange line to the a_T target box; Teacher Action a_T connects by an inner line to z_T. These reversed destinations conflict with the labeled video and action losses. A flow-matching box and JointLoRA update are also shown.
READING 53 · 3 ORIGINAL VISUALS

WAM-OPD

WAM-OPD trains actions under the student's own video plan using dense teacher labels, recovering capability in a two-task pilot while leaving target mismatch and stale training histories unresolved.

Read illustrated report →
Three-part architecture: an offline video teacher supplies a cached code; a causal history student predicts future, phase and horizon variables; two conditioning paths feed π0.5 beside the training losses.
READING 54 · 6 ORIGINAL VISUALS

ForeTime-VLA

ForeTime-VLA compresses privileged future-video supervision into causal history predictions that condition a flow-matching policy, improving reported conveyor grasping with a small measured latency cost.

Read illustrated report →
Three modality experts share a masked MoT attention layer; a tactile history branch accompanies a SAF encoder diagram and training/inference attention masks.
READING 55 · 5 ORIGINAL VISUALS

TacWAM

TacWAM combines mechanics-supervised tactile latents and recent contact history with a tri-modal generator whose attention lets tactile futures supervise learning while keeping them out of action inputs.

Read illustrated report →
Four-column comparison of direct cloning, world-action imitation and world-model policy learning, showing learned objects, population targets and decision principles.
READING 56 · 3 ORIGINAL VISUALS

Capability Separation

Future-action factorization can improve learning while preserving the ideal imitation target; stronger decisions require identified action effects and a deployment rule that compares their utility.

Read illustrated report →
History and masked future conditions enter an online encoder; an EMA target encoder supplies future targets. A shared predictor contains future, context and action streams, with the action branch added in Stage 2.
READING 57 · 6 ORIGINAL VISUALS

WA-JEPA

WA-JEPA couples latent-future flow generation with ego-trajectory prediction, improving simulated planning while leaving inference fidelity, comfort and deployment speed incompletely established.

Read illustrated report →
Three panels show DECOWAM's frozen WAN backbone and trainable conditioning modules, deployment without future frames, and a two-stage adaptation sequence; a legend distinguishes data, token, training, and conditioning arrows.
READING 58 · 6 ORIGINAL VISUALS

DECOWAM

Separating base, arm, and ego-motion information improves FastWAM's replay prediction and observed robot coordination with few Stage-2 trainable parameters, while retaining a large, slightly slower deployed model.

Read illustrated report →
Five observed surgical frames labeled I0 through I4 pass through a frozen SurgMotion ViT-L encoder into five token grids, with the original color legend retained.
READING 59 · 4 ORIGINAL VISUALS

Joint Surgical Visual–Trajectory Forecasting

A shared latent predictor improves surgical scene and tool-path forecasts by rolling forward in short chunks, while accumulated error limits its planning relevance.

Read illustrated report →
Four vertical diagrams compare LAPO, LAOF, CoMo, and CFD-AE. The first three reconstruct future frames; LAOF adds a flow decoder and CFD-AE reconstructs frame differences.
READING 60 · 6 ORIGINAL VISUALS

What Matters for Latent Actions

Video-derived latent actions can strengthen a robot policy's VLM initialization, but the best representation and integration strategy depend on the control task and training protocol.

Read illustrated report →
Training diagram on the left, Latent Evaluator and Rollout Gate in the center, and an inference loop on the right. The center prints “0 ≤” left of x_h toward Planner, conflicting with the stopping rule in the source caption and algorithm.
READING 61 · 6 ORIGINAL VISUALS

RISE

RISE learns when another imagined latent could improve a driving plan enough to justify its cost, trading exhaustive offline supervision for selective inference.

Read illustrated report →
Three-panel architecture: multimodal tokens and directed attention at left, contact-to-deformation-to-slip hierarchy at upper right, and candidate selection plus residual-based execution monitoring below.
READING 62 · 6 ORIGINAL VISUALS

HiTac-WAM

HiTac-WAM uses a directed contact–deformation–slip forecast to rank action candidates and monitor execution, improving three-task robot success while leaving generated-candidate forecast accuracy and slip calibration unresolved.

Read illustrated report →
Architecture with a blue video/action pathway on the left, a pink semantic/action pathway on the right, a purple Callosal Action Bridge between action streams, and green Cerebellar Intent Fusion before the action decoder. Insets show cross-attention gates and fusion self-attention.
READING 63 · 6 ORIGINAL VISUALS

BrainWAM

BrainWAM coordinates separate semantic and predictive action streams to improve simulated driving plans, while retaining an inference cost that still limits vehicle deployment.

Read illustrated report →
Three-panel architecture: alternative RGB-history encoders produce a prefix for a shared video–action DiT; modality decoders and a block-causal visibility matrix appear to the right.
READING 64 · 5 ORIGINAL VISUALS

WNM-3D

A geometry-aware history prefix improves joint video–action navigation on GN-Bench, while the curriculum shows that expert correction must precede effective reward refinement in the tested configurations.

Read illustrated report →
Two lanes connect multi-embodiment videos to offline flow-conditioned training, and an observed scene through Isaac Lab and projected action flow to online causal video rollout.
READING 65 · 6 ORIGINAL VISUALS

Hydra-0

Action flow translates visible motion into a shared video-conditioning interface, improving prediction across embodiments while leaving physical control dependent on calibrated grounding and a supervised action readout.

Read illustrated report →
Four-stage Teach-and-Grow diagram: demonstrations yield shared changes and blocks; an agent composes and checks execution; verified blocks and experience feed the next task.
READING 66 · 5 ORIGINAL VISUALS

Teach and Grow

TGL grows explicit, verified robot skills under fixed foundation weights, trading repeated agent-and-tool interaction for local capability updates whose broad scaling benefit remains untested.

Read illustrated report →
Three-panel architecture: a frozen Stage-JEPA branch adds guidance to WAM video tokens; video, action and understanding streams interact through joint attention; lower panels show stage-pair construction and predictor training.
READING 67 · 6 ORIGINAL VISUALS

StageWAM

A frozen next-stage latent predictor improves a joint video-action policy through gated conditioning, but the experiments support the extra pathway more strongly than task-specific stage learning.

Read illustrated report →
PhyAI hierarchy: scheduler points to stateful model runners containing stateless modeling modules; runners call shared Layers, which dispatch custom or library kernels to hardware. Pre/post-processing and offline quantization sit to the left.
READING 68 · 6 ORIGINAL VISUALS

PhyAI

PhyAI shares an execution runtime across existing VLA and WAM policies, accelerating model inference while leaving control deadlines and end-to-end rollout gains dependent on the surrounding system.

Read illustrated report →
An orange data synthesis, filtering, policy learning and simulation validation loop feeds a blue synthetic/real co-training block, followed by green physical deployment and evaluation.
READING 69 · 5 ORIGINAL VISUALS

RoboSynChallenge

RoboSynChallenge couples expandable simulated demonstrations with physical manipulation tests, but its initial baselines leave the value of co-training and several aggregate results unresolved.

Read illustrated report →
Grouped horizontal bars show end-effector drift under six future-cache interventions for four WAMs; color-matched success percentages and structural N/A entries appear at right.
READING 70 · 6 ORIGINAL VISUALS

RIFT

RIFT replaces iterative future-video generation with one learned cache prefill, retaining explicit future-conditioned actions at low deployment latency within the tested simulation settings.

Read illustrated report →
Two video-processing branches feed motion-difference and destination-distribution alignment losses; fire marks the trainable WAM branch and snowflakes the frozen teacher branch.
READING 71 · 6 ORIGINAL VISUALS

4D-WAM

Two training-only alignment losses transfer local motion and destination correspondence from a frozen trajectory teacher into existing WAMs, improving robustness while leaving deployment architecture unchanged.

Read illustrated report →
Video DiT and Action DiT blocks receive visual and action tokens; a purple dynamics register connects through an adapter to a latent-action target from a frozen encoder and IDM branch.
READING 72 · 6 ORIGINAL VISUALS

ForeWAM

ForeWAM reuses a single video-feature prefill to condition direct action denoising, achieving efficient benchmark control while leaving the causal meaning and broader robustness of its latent futures unresolved.

Read illustrated report →
Two panels show shared latent slots for current robot state/image and generated actions/future state/image, followed by video pretraining and joint fine-tuning.
READING 73 · 6 ORIGINAL VISUALS

Surgical WAM

A shared video-action diffusion model converts surgical video pretraining into better simulated manipulation under fixed action supervision, while joint future sampling and incomplete evaluation details leave efficiency and real-world transfer unresolved.

Read illustrated report →
Five method-family rows identify six strategies, their intervention mechanisms and whether retraining is required.
READING 74 · 6 ORIGINAL VISUALS

World Action Models in Real Time

Learning to continue an already committed action prefix balances execution quality and smoothness on this platform, but relies on temporal alignment and on the old trajectory remaining useful.

Read illustrated report →
FACT's shared causal diffusion transformer receives observation, noisy action, clean action, value and future-video tokens; separate panels show which losses remain active on successes and failures.
READING 75 · 6 ORIGINAL VISUALS

FACT

FACT uses a shared action-first transformer to learn failed consequences without imitating failed actions, improving control while leaving consequence scoring an optional computation cost.

Read illustrated report →
Separate action and video branches exchange attention representations. Predicted and recorded future frames enter a frozen geometric teacher, whose feature and depth losses send gradients back to the backbone. A three-group attention mask protects condition tokens.
READING 76 · 6 ORIGINAL VISUALS

4D-WAM

4D-WAM improves joint video-and-trajectory prediction by supervising generated futures with a frozen geometric teacher and emphasizing early decision timesteps, while keeping that teacher outside inference.

Read illustrated report →
Two training branches share a MoT predictor: inverse dynamics predicts action velocity, forward dynamics predicts future latents against a stop-gradient EMA target; adjacent grids show allowed and blocked attention.
READING 77 · 6 ORIGINAL VISUALS

SLIM-0.5B

SLIM uses bidirectional latent prediction to train a compact action policy, retaining predictive structure at deployment while avoiding explicit future generation.

Read illustrated report →
Training architecture with a bottom VLM–World Adapter–action-head path and an upper video denoiser conditioned by world tokens and an encoded edge map.
READING 78 · 6 ORIGINAL VISUALS

World Tokens

Training a video denoiser through the policy’s mandatory token interface improves manipulation while moving video-model cost out of deployment and into training.

Read illustrated report →
Architecture showing task inputs, the VLM Task Manager and evidence memory above a deterministic plan compiler, with WAM execution and progress/recovery feedback below.
READING 79 · 6 ORIGINAL VISUALS

HarnessWAM

HarnessWAM improves long-task execution by combining evidence memory, constrained skill compilation, and event-triggered verification, while remaining limited by the executor’s validated skills and the reliability of semantic judgments.

Read illustrated report →
TempoWAM flow diagram: a WAM generates actions, an initial prefix executes, and RPM plus AEP accepts another prefix or discards the remaining chunk. Easy and hard stages have different reuse patterns.
READING 80 · 6 ORIGINAL VISUALS

TempoWAM

TempoWAM uses predicted task progress to spend replanning calls where action chunks become unreliable, trading computation against success through a calibrated execution rule.

Read illustrated report →
Two V-JEPA encoding paths meet a shared Qwen predictor: a paired-image teacher supervises visual-position predictions, while separate orange action states feed a DiT action expert.
READING 81 · 6 ORIGINAL VISUALS

JEPA-WAM

Dense current–future embedding supervision improves action learning through a shared predictor, while deployment keeps the action pathway and omits explicit transition prediction.

Read illustrated report →
Training and inference diagram with frozen teacher and IDM, video and action tokens, four supervision terms, and a student-only inference panel.
READING 82 · 6 ORIGINAL VISUALS

Vid2WAM

Vid2WAM moves video generation and inverse dynamics into offline supervision, improving a student's action policy while leaving difficult held-out tasks far from solved.

Read illustrated report →
Three-stage architecture linking a whole-body VLM, V-JEPA features, dual video/action queries, a 0.45B action DiT and SONIC; the right panel details attention direction and positional encodings.
READING 83 · 6 ORIGINAL VISUALS

ω-0

ω-0 uses future visual supervision to improve directly denoised humanoid action latents, with controller grounding and action continuity carrying much of the practical burden.

Read illustrated report →
Eight model diagrams: four VLA structures above four WAM structures, with purple inputs, green intermediate representations, orange outputs and gray model blocks.
READING 84 · 4 ORIGINAL VISUALS

Embodied.cpp

Embodied.cpp unifies deployment around a C++ runtime, with reported latency and memory gains whose control-quality cost depends on model and precision.

Read illustrated report →
Five labeled layers ascend from General Data through Simulation Data, Ego Data, and UMI Data to Robot Data, with example interaction images underneath.
READING 85 · 6 ORIGINAL VISUALS

Embodied Data Pyramid

Scalable embodied learning needs complementary data sources whose supervision is aligned to physical execution; the survey maps those tradeoffs without establishing a universally best mixture.

Read illustrated report →
Two camera observations of one state enter the same WAM denoiser. Blue action, future-proprioception and value outputs are linked across views; the orange future-image link is crossed out.
READING 86 · 6 ORIGINAL VISUALS

SCVC

Selective agreement across camera views improves simulated control beyond trained viewpoint ranges, while leaving future images view-dependent and offering no interpolation benefit.

Read illustrated report →
Three-branch PILOT diagram with a Wan2.2 world-model context path, Action-Perceiver state/query/action tokens, asymmetric attention matrices and a Causal Dynamics Engine trained against future features.
READING 87 · 6 ORIGINAL VISUALS

PILOT

PILOT uses training-time future-feature prediction to shape protected motion tokens for faster action generation, but its reported gains come with unresolved implementation and evaluation inconsistencies.

Read illustrated report →
Video DiT and Action DiT with separate flow-matching losses, a frozen DINOv3 future-frame teacher, semantic query tokens and shared positional encodings.
READING 88 · 6 ORIGINAL VISUALS

Robust-WAM

Robust-WAM adds future semantic supervision to an existing WAM's action stream, improving tested OOD success while retaining the pretrained video formulation and extra query tokens at inference.

Read illustrated report →
Architecture diagram with image and text inputs, Wan blocks 1 through 30, intermediate trajectory readouts, a quality scorer, early-exit decision and future-scene training branch.
READING 89 · 5 ORIGINAL VISUALS

Adaptive-WAM

Intermediate video features can support strong driving plans, while a learned quality verifier trades backbone depth for latency without a safety guarantee.

Read illustrated report →
Three-panel architecture showing frozen visual/text encoders, world and action transformers, an attention-mask inset, a training-only foresight branch and three softly weighted action experts.
READING 90 · 6 ORIGINAL VISUALS

MobileWAM

MobileWAM trains current-frame representations with recurrent future supervision and softly mixes motion experts, improving mobile control while eliminating future generation—but retaining the large backbone—at deployment.

Read illustrated report →
Architecture diagram with VAE-encoded current and history-flow frames, text cross-attention, a kinematic-conditioned action stream, joint attention, two decoders and three training stages.
READING 91 · 6 ORIGINAL VISUALS

DynamicWAM

DynamicWAM couples spatial flow history with image-plane motion statistics to improve interception, while cached video inference and asynchronous execution address a separate source of delay.

Read illustrated report →
Training schematic with paired VideoDiT and ActionDiT, RGB/flow channel concatenation, depth and DINO prediction heads, and an enlarged block with gated residual branches.
READING 92 · 5 ORIGINAL VISUALS

DreamWAM

DreamWAM improves robustness by learning several views of the future through representation-specific training routes, while keeping auxiliary teachers and predictions out of RGB-only deployment.

Read illustrated report →
Video and action branches with green interval-fusion blocks; a side panel contrasts repeated joint updates with one cached video pass and repeated sparse action updates.
READING 93 · 6 ORIGINAL VISUALS

Faster-WAM

A video expert can supply useful future-aware context from one noisy pass, while sparse access and depth-wise fusion reduce its repeated cost during action generation.

Read illustrated report →
LiLa-WAM diagram: frozen DINOv3 current-frame features pass through fusion and a query adapter into an expert containing visual, VTT, proprioceptive and action tokens; foresight outputs feed a training-only decoder supervised by future-frame features.
READING 94 · 6 ORIGINAL VISUALS

LiLa-WAM

LiLa-WAM makes future prediction an auxiliary teacher inside a compact action transformer, gaining control performance while relying on a fixed demonstration-derived task cue.

Read illustrated report →
Two panels show training with six token types and three losses, followed by joint refinement of noisy images and trajectories conditioned on history and goal.
READING 95 · 6 ORIGINAL VISUALS

UniNav

UniNav learns video, geometry and waypoints in one transformer, while a separately trained Fast variant trades inference-time visual foresight for lower waypoint-prediction latency.

Read illustrated report →
Four-column diagram connects a frozen action–future candidate pool, active coordination contracts, ensemble verification and conjunctive gates to override, preserve or abstain, with a separate offline oracle audit.
READING 96 · 6 ORIGINAL VISUALS

CoWAM

CoWAM converts predicted bimanual futures into conservative action overrides through coordination contracts, trading some available rescues for fewer unsupported interventions.

Read illustrated report →
A video transformer hub sends keys and values through KV-Fusion docking interfaces to separate task heads; each head has its own projections and layer aggregation.
READING 97 · 5 ORIGINAL VISUALS

Faster-WAM

Faster-WAM lets a one-layer action head read all 30 video-transformer layers, reducing reported inference latency while trading some RoboTwin accuracy for a stronger LIBERO-Plus result.

Read illustrated report →
Main and wrist images, language and dynamics tokens enter a VLM; full context conditions an action expert. Upper branches show frozen VGGT supervision and an action-conditioned predictor aligned with an EMA future target.
READING 98 · 6 ORIGINAL VISUALS

SG-WAM

SG-WAM uses geometry-aligned policy tokens and action-conditioned future prediction to improve a compact manipulation policy, while keeping the predictive branches out of deployment.

Read illustrated report →
Video DiT in the center receives a current frame and instruction; its hidden states connect to a future-target Grounding DiT on the left and an action expert on the right.
READING 99 · 6 ORIGINAL VISUALS

EndoWAM

EndoWAM trains a video-model representation to reconstruct future target regions, then uses that representation for fast discrete endoscopic control while omitting the grounding branch at deployment.

Read illustrated report →
Two modality-specific Transformer experts connect through mixed attention; training and inference matrices show allowed query-to-key access for current frames, future video, clean actions and future actions.
READING 100 · 6 ORIGINAL VISUALS

SelfWAM

SelfWAM uses selectively action-conditioned RGB and robot-mask prediction to improve a direct policy while keeping future generation out of the deployed control loop.

Read illustrated report →
Blue video and orange action streams pass upward through separate encoders, shared video-action attention, separate feed-forward networks, and output decoders; the action output labels control points b3 through b7.
READING 101 · 6 ORIGINAL VISUALS

FlowPilot

FlowPilot jointly denoises future-depth latents and a compact polynomial trajectory to improve agile navigation, while its onboard timing and feasibility claims require narrower interpretation than its headline suggests.

Read illustrated report →
Three input paths—current RGB/tactile observations, 8-D robot state and a successful future suffix—feed ACT. An inset contrasts asymmetric and isolated modality-reading matrices and four phase anchors per modality.
READING 102 · 4 ORIGINAL VISUALS

OVTF / AFM

With successful reference futures fixed, selective phase-local visual-to-tactile routing improves simulated action success, while leaving learned-future reliability and deployability untested.

Read illustrated report →
Four panels connect multimodal dependency, controlled visual/audio omission, ARR and LAD diagnosis, and an MSA module with SVD bases, projection norms, sorting and hidden-state correction.
READING 103 · 5 ORIGINAL VISUALS

MSA and Causal Modality Sensitivity

Modality subspace steering makes Omni-LLM outputs more responsive to missing sensory inputs, but the reported gains in sensitivity do not consistently preserve answer accuracy.

Read illustrated report →
Blue diagrams show separate state and action Jacobian corrections; green diagrams show a joint block Jacobian carrying a new state residual into both state and action coordinates.
READING 104 · 6 ORIGINAL VISUALS

FBFM

FBFM uses measured state discrepancies to steer a frozen WAM's active flow solver, improving some control outcomes while making timing, residual scale and latent compatibility critical.

Read illustrated report →
Flow diagram from query through Medium and the prediction-interface router to Medium or Full action selection, with a separate exact-reset outcome and paired-label branch.
READING 105 · 6 ORIGINAL VISUALS

Paired World-Model Cascade Evaluation

A Medium-derived prediction interface can improve selective physical decisions enough to pay for sequential overhead, while still taking longer than fixed prediction policies.

Read illustrated report →
Three rows of predicted futures drift from shifted visual conditions toward LIBERO appearance; below, clean and shifted frame triplets accompany DINOv3 and Wan-VAE cosine-similarity distributions.
READING 106 · 6 ORIGINAL VISUALS

ST-WAM

ST-WAM combines VAE dynamics, DINO future supervision and current-conditioned history to improve robustness under visual shifts, while paying additional inference cost without generating futures at deployment.

Read illustrated report →
Three calibration panels show compatible-channel pooling, joint video/action saliency and fixed-count replay auditing above video/action DiT blocks and action chunks.
READING 107 · 6 ORIGINAL VISUALS

QuantWAMs

QuantWAMs spends precision where compatible coordinates, joint video–action gradients and controlled rollout replay support it, trading additional calibration work for lower targeted-block resource use.

Read illustrated report →
Architecture diagram connecting text, noisy-video and skeleton embeddings to autoregressive DiT chunks, with metric rotary encoding and separate anchor/recent scene-memory slots.
READING 108 · 6 ORIGINAL VISUALS

EgoGenesis

A fixed scene anchor, refreshed generated memory and metric action attention make egocentric synthesis more useful for robot training, while still requiring supplied trajectories and a separate control model.

Read illustrated report →
Two panels contrast noisy semantic and action training streams with an inference pipeline that encodes the current observation into a semantic cache for the Action DiT.
READING 109 · 6 ORIGINAL VISUALS

LeapBot-WA

Future semantic guidance trains the action policy, while a static observation cache removes future rollout at deployment; the exact training recipe remains inconsistent in the supplied v2.

Read illustrated report →
DC-WAM diagram with dense temporal differences, sparse weighted flow error, CoTracker3 motion-map construction, and DynaRoute bias entering a two-branch MoT attention block.
READING 110 · 6 ORIGINAL VISUALS

DC-WAM

DC-WAM concentrates video supervision and attention on motion to improve manipulation robustness, while using a single visual-cache prefill for action-only denoising instead of repeatedly generating future video.

Read illustrated report →
Six camera views, hand-held glove grippers and marker cubes, head and hand trajectories, handwriting reconstruction, and a wearer with labeled capture components.
READING 111 · 6 ORIGINAL VISUALS

HiFi-UMI

A wearable capture-and-curation pipeline supplies deployable task-specific supervision, with near-teleoperation aggregate success obtained from substantially more robot-free demonstrations.

Read illustrated report →
Architecture connecting image, instruction, estimated pose and robot state to Understanding, Memory, Evolution and heterogeneous-policy modules; output subtasks are separated by Memory Bridges.
READING 112 · 6 ORIGINAL VISUALS

RoboHarness

RoboHarness improves long-horizon execution by selecting complementary controllers and preparing their handoff states, while depending on accumulated memory and adaptive orchestration.

Read illustrated report →
Green video, blue tactile and cream action experts share an attention bar; a separate observed-tactile encoder feeds cross-attention near execution.
READING 113 · 6 ORIGINAL VISUALS

N₀-TWAM

N₀-TWAM predicts scene and contact before acting, then adds current-touch feedback; its average manipulation gains come with sensor-dependent representations and task-specific deployment choices.

Read illustrated report →
Three modules connect video geometry tokens to future geometry aggregation and a trajectory planner. Ego status enters the future and action modules; the planner aggregates present geometry before future tokens. Reconstruction and depth examples appear at right.
READING 114 · 6 ORIGINAL VISUALS

GeoWorldAD

GeoWorldAD refines driving trajectories with ego-aligned present geometry and latent future geometry, improving simulated progress while leaving the causal contribution of added training and the limits of forecasting unresolved.

Read illustrated report →
Four camera-specific DiT streams with view/time conditioning, causal self-attention, a shared cross-view connection, text cross-attention and per-view denoised outputs.
READING 115 · 6 ORIGINAL VISUALS

OmniDreams

OmniDreams turns a video backbone into a responsive camera simulator, trading substantial compute and chunk-level feedback for controllable visual generation.

Read illustrated report →
Stages 1a and 1b adapt InternVL3 and distill geometry, semantics, and dynamics; Stage 2 routes frozen perception and video features through queries, a shared Transformer, future prediction, and an action head.
READING 116 · 6 ORIGINAL VISUALS

PerceptDrive

PerceptDrive preserves distinct frozen perception priors and routes their influence before future-conditioned trajectory generation, trading evaluator-specific training for inference without candidate selection.

Read illustrated report →
Architecture with T5 and VLM branches, three event-memory views, query-based retrieval and gating, causal video-action DiT blocks, and a real-observation feedback arrow.
READING 117 · 6 ORIGINAL VISUALS

WorldScape Policy 2.0

WorldScape Policy 2.0 turns event captions and retrieved task history into conditions for a shared video-action model, improving measured control while leaving clean-only transfer and implementation completeness as important limits.

Read illustrated report →
Left-to-right diagram of primary and wrist images entering a WAM, an action–future cosine gate, optional candidate sampling and wrist-to-primary depth reprojection, ending in action execution.
READING 118 · 6 ORIGINAL VISUALS

Gated GeoBoN

Gated GeoBoN spends extra WAM inference compute when action and imagined motion disagree, then selects geometrically consistent futures, trading some always-on success gain for lower average latency.

Read illustrated report →
Two panels show action-centered joint training and efficient inference. Visual and action experts exchange information, while dashed inference arrows mark optional future processing and a KV cache appears below the action expert.
READING 119 · 6 ORIGINAL VISUALS

GigaWorld-Policy-0.5

Future-dynamics supervision and a smaller action expert support stronger robot policies with an 85 ms C++ inference path, although the study does not isolate every source of the improvement.

Read illustrated report →
Three-part diagram of clean WAM outputs, bounded online observation perturbation through a frozen model, and divergent imagined versus executed paths; a lower strip shows paired query search and its legend.
READING 120 · 6 ORIGINAL VISUALS

BadWAM

BadWAM finds observation perturbations that degrade closed-loop control while preserving comparatively similar predicted futures, exposing the limits of imagination as a safety signal.

Read illustrated report →
A shared DiT receives state, reference-image, action and future-image tokens. A query-by-key attention matrix blocks future-image keys from action queries. A dashed box marks future-video components disabled at inference.
READING 121 · 6 ORIGINAL VISUALS

AeroAct

AeroAct uses one Transformer to learn trajectory actions and their visual consequences, then removes future-video prediction for flight; temporal history is strongly supported in simulation, while physical evidence remains a short indoor demonstration.

Read illustrated report →
Twelve three-dimensional activation plots compare Cosmos-Policy and DiT4DiT under noise and camera perturbations, showing positive and negative points, SVM planes, block labels and hinge losses.
READING 122 · 6 ORIGINAL VISUALS

WA-LQR

Feedback steering can improve a pretrained WAM's robustness when nuisance features occupy a transferable low-dimensional subspace, but its gains depend on task geometry and calibration.

Read illustrated report →
RGB and flow branches enter separate patch embeddings and a shared DiT; a separate action expert follows. Lower panels distinguish two training stages and policy/world-model inference.
READING 123 · 5 ORIGINAL VISUALS

FlowWAM

Encoding motion as a flow video lets one visual generator support control and prediction, while a separate action expert and carefully constructed flow targets remain essential.

Read illustrated report →
Two panels separate a WAM consequence-prediction prototype from a brain–harness execution loop that returns observed outcomes through Trace Cards and evaluation/admission.
READING 124 · 6 ORIGINAL VISUALS

From WAMs to Embodied Brains

The roadmap proposes reusable physical reasoning through explicit consequence, intent and execution contracts, while leaving transfer benefits and the final model architecture to empirical testing.

Read illustrated report →
Two training branches: frozen visual features enter a vision inverse/forward dynamics pathway; action chunks enter an action encoder and decoder. Visual latents supply global context to action reconstruction, with gradient feedback toward the visual representation.
READING 125 · 6 ORIGINAL VISUALS

Lumo-2

Lumo-2 uses progressive multimodal alignment to make latent dynamics useful for fast action generation, while its experiments leave the separate effects of supervision, history and online prediction unresolved.

Read illustrated report →
AFP diagram: RGB frames feed MobileNetV3 and a dense feature path; learned region queries cross-attend to encoded features, receive temporal modeling and CLIP/FiLM language conditioning, and combine with dense features to produce a mask used for auxiliary policy training.
READING 126 · 6 ORIGINAL VISUALS

AFP

AFP transfers task-relevance masks into a policy’s attention during fine-tuning, improving reported distractor robustness without an inference-time mask predictor, at the cost of extra annotation and training machinery.

Read illustrated report →
Three panels show bank-query visual summarization and memory conditioning of Wan-DiT, auxiliary training supervision, and mass-weighted merging of historical events with protected anchor and latest slots.
READING 127 · 6 ORIGINAL VISUALS

DiM-WAM

DiM-WAM improves temporally dependent manipulation by conditioning joint video/action denoising on bounded observation memory, while relying on coarse progress supervision whose transfer beyond the tested tasks remains unresolved.

Read illustrated report →
Execution path from observation s through a latent policy and frozen generative policy to the environment; expert corrections feed action inversion and supervised latent-policy updates.
READING 128 · 5 ORIGINAL VISUALS

FlowDAgger

FlowDAgger turns expert actions into supervision for a frozen robot policy's noise input, improving adaptation while remaining limited by the behaviors that generator can express.

Read illustrated report →
Task, RGB/depth and robot state feed an agentic planner connected to a primitive library, an action interface, and separate task-specific and global memories.
READING 129 · 5 ORIGINAL VISUALS

Harness VLA

Harness VLA improves simulated manipulation by combining memory-guided staging and recovery with a frozen contact policy, at the cost of reference-task exploration and a larger execution system.

Read illustrated report →
Cosmos video flow matching receives a language instruction and an anchored observation; yellow intermediate features and a video-noise signal pass through feature projection to a Gemma action head with state and noisy action tokens.
READING 130 · 6 ORIGINAL VISUALS

Temporal Ratio

Schedule stronger video conditioning when action attention shifts toward imagined futures, while retaining current-observation grounding for precise manipulation.

Read illustrated report →
Three-part pipeline: coupled action/video diffusion experts, a TTT residual with robot and human token projections, and deployment human videos updating fast weights inside a frozen WAM.
READING 131 · 6 ORIGINAL VISUALS

WAM-TTT

Human-only fast-weight adaptation can steer coupled video/action prediction, but its transfer depends on a previously learned human–robot interface and incompletely specified evaluation separation.

Read illustrated report →
One transformer receives ego vision, proprioception, robot wrist vision and learned action/future tokens. Two heads read its features; red dashed loss paths return to the trunk. Three right-hand panels illustrate alternative world heads.
READING 132 · 5 ORIGINAL VISUALS

EgoWAM

Predicting semantic features or stabilized motion during training improves transfer from human demonstrations, while deployment retains only the action policy.

Read illustrated report →
A video-action backbone sends intermediate features to four depth extraction blocks. Two query/key matrices show allowed attention among history video, future video, actions and depth registers.
READING 133 · 6 ORIGINAL VISUALS

WAM4D

WAM4D uses future-depth supervision to train causal video-action features, then removes the geometry branch for efficient control, with gains that depend on the evaluation setting.

Read illustrated report →
Training diagram with video, action and auxiliary 4D geometry experts; frozen VGGT supplies geometry, while the inference panel retains only video and action experts.
READING 134 · 6 ORIGINAL VISUALS

MECo-WAM

MECo-WAM uses temporary geometric guidance and action-weighted relational supervision during training to improve manipulation while retaining the base video-action inference graph.

Read illustrated report →
UNIVERSE pipeline with text, history images and velocity encoders, shared diffusion transformer, masked context/video/action attention matrices, and three rollout modes.
READING 135 · 6 ORIGINAL VISUALS

UNIVERSE

UNIVERSE uses video supervision to train a shared trajectory denoiser, while a visibility mask removes the need for future-video generation during planning.

Read illustrated report →
Three-part overview: real-world training data on the left, optional language subtask planner in the center, and a Wan2.2-TI2V-5B video model with action expert on the right. Arrows carry observation, instruction and subtask conditioning; insets compare inference latency.
READING 136 · 6 ORIGINAL VISUALS

DSWAM

DSWAM pairs an optional language planner with a video-co-trained executor that generates action chunks directly, trading explicit future-video inference for lower query latency.

Read illustrated report →
Architecture schematic with video features in yellow, frame-level latent actions in orange, manipulation in blue and mobility in green, connected through shared attention and two training stages.
READING 137 · 6 ORIGINAL VISUALS

ABot-M0.5

ABot-M0.5 uses a shared, specialized video-to-motion-to-control architecture and dreamed-future training to improve robot success, with unresolved sampling details and uneven generalization.

Read illustrated report →
Five-stage table comparing single-environment and batched policy initialization, observation updates, action prediction, execution and evaluation roles.
READING 138 · 6 ORIGINAL VISUALS

RoboDojo

RoboDojo makes manipulation failures easier to measure through complementary simulation and physical tests, while standardizing interfaces and reducing evaluation cost.

Read illustrated report →
Architecture diagram with robot state, task prompt, and three RGB views entering a frozen WA backbone. Reference-action and visual tokens feed a trainable actor and critic. Both use self-attention then cross-attention; the critic additionally receives candidate actions and outputs twin Q estimates.
READING 139 · 5 ORIGINAL VISUALS

HALO-WA

A frozen world-action model supplies reference actions and visual memory to an online actor-critic, improving precision manipulation while depending on task-specific preparation and a competent base policy.

Read illustrated report →
Architecture with frozen VAE and vision-language encoders, robot-state conditioning, separate video/action paths inside mixed attention, training and inference masks, and video/action outputs.
READING 140 · 6 ORIGINAL VISUALS

Kairos

Kairos trains video and action transformers together, retaining an action-only inference path, but its benchmark gains do not yet establish closed-loop regret reduction.

Read illustrated report →
Three visual, tactile and action expert streams, training and inference attention masks, and a contact-gated auxiliary attention loss with stop-gradient keys.
READING 141 · 5 ORIGINAL VISUALS

VT-WAM

VT-WAM predicts tactile evolution alongside actions and guides contact-phase attention, retaining tactile computation while dropping future visual prediction during control.

Read illustrated report →
Architecture with an orange frozen-teacher/cache branch above a Florence-2 encoder, three prior predictors, and a WorldBridge action transformer ending in an action chunk.
READING 142 · 6 ORIGINAL VISUALS

Bridge-WA

Bridge-WA trades a deployment-time future generator for compact predicted outcome, change and motion priors, whose usefulness depends on how the action transformer reads them.

Read illustrated report →
Six manipulation images: simulator banana, drawer and brick scenes on the left, and corresponding real-world object scenes on the right; group labels are retained.
READING 143 · 4 ORIGINAL VISUALS

WAM Sim-to-Real Transfer

Automated, randomized simulation demonstrations give a pretrained world-action policy measurable real-robot success, but this early study leaves the causes of transfer and its reliability largely unisolated.

Read illustrated report →
RGB history and current observation feed frozen visual and spatial encoders; their projected features merge before a shared FutureNav backbone with four labeled output branches.
READING 144 · 6 ORIGINAL VISUALS

FutureNav

FutureNav strengthens a shared navigation policy with spatial features and auxiliary transition learning, while leaving future-prediction heads inactive during default action decoding.

Read illustrated report →
Start and goal RGB images and DA3-estimated depths enter VAE encoders. A single DiT transforms visual and action tokens, with branches to the VAE decoder, VGAR, and action decoder. TSR is shown below the trajectory output.
READING 145 · 6 ORIGINAL VISUALS

SWAM

Joint RGB-D and action denoising improves offline navigation planning, while endpoint regularization and visual refinement bring distinct benefits and unresolved execution limits.

Read illustrated report →
A recurrent WAM diagram with a compact policy block on the left and blue initialization followed by green generation steps on the right. Predicted observations feed later policy inputs; action outputs are collected into a pseudo-trajectory.
READING 146 · 6 ORIGINAL VISUALS

REGEN

A world-action policy can rehearse old skills through recurrent synthetic trajectories, but retention depends on whether its imagined observations and actions remain coherent.

Read illustrated report →
Multi-view images branch through frozen DA3 and spatial-temporal VAEs into camera geometry, latent tokens and relation supervision for a trainable multi-view DiT.
READING 147 · 6 ORIGINAL VISUALS

PAIWorld

PAIWorld couples camera-aware attention with frozen-teacher relation distillation to improve multi-view video generation, while leaving the connection to executed robot success unmeasured.

Read illustrated report →
Two transformer streams share an attention operation: blue autoregressive vision/language tokens and orange diffusion vision/audio/action tokens. The mask blocks autoregressive queries from diffusion keys, while diffusion queries attend both streams; the legend distinguishes modalities and noisy tokens.
READING 148 · 6 ORIGINAL VISUALS

Cosmos 3

Cosmos 3 jointly denoises future video and actions inside a broader multimodal architecture, gaining transferable control priors while leaving compute-matched synergy and causal rollout reliability incompletely established.

Read illustrated report →
GIC architecture with observations ascending from the universe through a belief encoder and configurator, simulated trajectories above, and actions descending to an illustrated aircraft sequence. Blue policy, orange world-model and green critic operations have a legend.
READING 149 · 6 ORIGINAL VISUALS

GIC: Critique of Agent Model

GIC proposes persistent goals, adaptive identity and learned control of simulation while keeping the dynamics model independently trained; its benefits remain conditional theoretical claims awaiting full-system evaluation.

Read illustrated report →
Observation images, instruction and robot state feed a masked action sequence; step log-probabilities and environmental rewards connect to a value head and PPO policy update.
READING 150 · 6 ORIGINAL VISUALS

dVLA-RL

Scoring newly unmasked action tokens across denoising steps enables PPO refinement of a discrete VLA, while task-specific horizons trade computation against action quality.

Read illustrated report →
Left-to-right diagram of keyed sampling noise entering a VLA or WAM, a partial executed-action window feeding MAP recovery, candidate-key comparison, and rollout-score aggregation for verification or identification.
READING 151 · 6 ORIGINAL VISUALS

Keyed Latent Provenance

A keyed sampler leaves recoverable provenance in executed actions, but the evidence depends on calibration and continued use of that sampling process.

Read illustrated report →
An observation passes through encoder E and world model W to an adversarial imagined future, which branches to an oracle and a denoiser detector. A dashed red gradient path returns to the observation.
READING 152 · 6 ORIGINAL VISUALS

Attacking the Trusted Imagination

Observation attacks can corrupt imagined futures and impair a simulated MPC, but latent damage, detector evasion and imagination-specific task failure require separate evidence.

Read illustrated report →
Three-stage GAM diagram: shallow geometric encoding, a language-conditioned causal predictor, and deep geometric/action decoding; adjacent attention matrix has a fully enabled language-query row.
READING 153 · 6 ORIGINAL VISUALS

GAM

GAM makes geometric prediction part of action decoding inside a shared backbone, improving reported camera robustness while leaving the causal role of future supervision and several implementation details unresolved.

Read illustrated report →
Video and action-value transformer branches share attention. A colored block mask restricts modality access; side panels show two training stages and progress-triggered rollback.
READING 154 · 6 ORIGINAL VISUALS

MV-WAM

MV-WAM improves visually shifted manipulation by giving coupled video and action experts different prediction targets, while its value-guided recovery remains operationally underspecified.

Read illustrated report →
Definition diagram comparing direct VLA action prediction, world-model consequence prediction, and WAM future/action coupling; lower panels show prediction before action, consequence scoring and joint prediction.
READING 155 · 5 ORIGINAL VISUALS

World Action Models: A Survey

The survey organizes WAMs by how predicted futures influence control, arguing that selective imagination can preserve useful dynamics while reducing the cost of acting.

Read illustrated report →
Two coupled transformer branches: orange video tokens and hybrid memory on the left, blue noisy action tokens and action decoding on the right, with text and proprioceptive conditioning.
READING 156 · 6 ORIGINAL VISUALS

MemoryWAM

MemoryWAM keeps detailed initial and recent observations alongside per-frame gist memory, improving memory-dependent control while reducing the cost of a still-growing historical cache.

Read illustrated report →
EventVLA diagram with visual anchors and event images above a unified VLA, action-token branches to action and keyframe heads, and a future delayed-commit loop returning an observed frame to memory.
READING 157 · 6 ORIGINAL VISUALS

EventVLA

EventVLA forecasts when to save actual camera observations, improving transient-evidence manipulation through sparse memory while adding latency and retaining a fixed-capacity bottleneck.

Read illustrated report →
Instruction tokens and an encoded observation feed an image-editing backbone; a rightward KV arrow connects it to an action expert receiving robot state and action noise. Snowflakes and flames mark frozen and tunable modules.
READING 158 · 6 ORIGINAL VISUALS

ImageWAM

ImageWAM learns robot actions from image-editing caches, avoiding decoded future videos while making its strongest efficiency and robustness claims dependent on the comparison setup.

Read illustrated report →
Side-by-side training diagrams: actor-only sparse-reward learning on the left; world-model prediction, actor execution, reconstruction feedback and success-filtered video SFT with KL regularization on the right.
READING 159 · 5 ORIGINAL VISUALS

WAM-RL

WAM-RL improves a video world model and its action translator together, using reconstruction rewards and constrained video adaptation to balance execution fidelity against latent-feature drift.

Read illustrated report →
Two specialist transformer branches receive encoded video and action tokens; text and ego-state conditioning enters from the left. Training and inference insets distinguish current-frame, future-frame and action paths.
READING 160 · 6 ORIGINAL VISUALS

Metis

Metis uses action-conditioned video prediction to train a trajectory policy, then removes future-video synthesis from inference to reduce latency.

Read illustrated report →
Two training stages: frozen DINO features feed an inverse-dynamics encoder and LaWM decoder; a vision-language policy then predicts a latent action whose decoded subgoal conditions an Alternate-DiT action expert.
READING 161 · 6 ORIGINAL VISUALS

LaWAM

LaWAM makes predicted DINO features an explicit action-generation input, exchanging pixel-level future synthesis for a compact dynamics interface whose evidence is strongest in stable-camera manipulation.

Read illustrated report →
Current robot observation queries a trajectory database; retrieved observations and actions condition ReCAP alongside the current frame, producing a target action chunk and future observation.
READING 162 · 6 ORIGINAL VISUALS

ReCAP

ReCAP adds tasks to a frozen world-action policy through retrieved trajectories, while learned residuals adapt their motions to a compatible target embodiment.

Read illustrated report →
Episode images enter a Perceiver compressor. Purple memory tokens condition either an SVD UNet or Cosmos DiT backbone, join video-derived action conditioning, and enter a three-stream completion gate.
READING 163 · 5 ORIGINAL VISUALS

MemoryVAM

MemoryVAM makes episode history available to both video prediction and action decoding, improving memory-dependent manipulation while retaining a history cache whose cost grows with episode length.

Read illustrated report →
A video autoencoder aligned with a visual foundation model sits above inverse and forward dynamics blocks; transport K and residual delta reconstruct the next latent.
READING 164 · 6 ORIGINAL VISUALS

RepWAM

RepWAM makes visual states and latent actions share a semantic representation, improving the reported control ablations while leaving the robot adaptation interface incompletely specified.

Read illustrated report →
Architecture with visual history, three goal types and motion history feeding two contextual streams, conditioning tokens and a shared action/latent generator.
READING 165 · 6 ORIGINAL VISUALS

WAM-Nav

WAM-Nav couples a long action trajectory to one-step latent visual foresight inside a shared generator, improving reported navigation performance while retaining task-balancing and embodiment limits.

Read illustrated report →
Three panels show joint RGB-mask/action training, permitted attention between token blocks, and cached visual context feeding action denoising at inference.
READING 166 · 6 ORIGINAL VISUALS

MaskWAM

Predicting a target's future mask makes first-frame visual prompting useful for joint action generation, at the cost of mask annotation and segmentation dependence.

Read illustrated report →
A Cosmos Predict2 2B block takes blank, state, goal-image and current-image latent frames below, and produces action, future-state, two future-image and goal-progress frames above.
READING 167 · 6 ORIGINAL VISUALS

NavWAM

NavWAM turns a video diffusion transformer into an action-producing navigation policy, with evidence that future-image supervision helps offline accuracy and limited-scale evidence of physical deployment.

Read illustrated report →
A left-to-right diagram connects the current camera frame to a frozen VAE encoder, action embeddings and a conditioned DiT, with a blue present-latent bypass into an addition node before frozen decoding.
READING 168 · 6 ORIGINAL VISUALS

DiT for AV Scene Prediction

Calibrated latent diffusion improves the appearance distribution of compact driving predictions, while a separate re-anchored jump model improves coarse motion at the cost of accumulating blur.

Read illustrated report →
Two-panel architecture: a video DiT and force encoder feed an action–force DiT at 1 Hz; intervention-trained residual and gate heads use predicted and measured forces to correct robot commands at 10 Hz.
READING 169 · 6 ORIGINAL VISUALS

FAWAM

FAWAM turns predicted wrist wrenches into references for fast, gated action correction, improving contact-rich task completion while requiring additional intervention data and careful gate selection.

Read illustrated report →
EWAM schematic with observation, robot state and language inputs, a gray world/action backbone, blue memory and correction modules, environment feedback, and a red-bordered filtering and learning loop.
READING 170 · 6 ORIGINAL VISUALS

EWAM

EWAM trades extra inference work for more efficient simulated execution by adding memory, anomaly-aware correction and filtered online adaptation to a frozen WAM.

Read illustrated report →
A blue VLM branch and an orange frozen WAM branch receive observations and instructions. Future-scene tokens pass through a dynamics encoder into Latent Steering; anticipated actions pass through an action encoder into the action head, which also receives noise.
READING 171 · 6 ORIGINAL VISUALS

World Pilot

A frozen world-action model supplies future-scene features and a soft motion hint to a separate VLA, improving reported OOD success at the cost of an extra model pass per decision.

Read illustrated report →
Baseline future frames and action-attention overlays beside current/generated views and zero/mean hidden-state intervention heatmaps; the right-hand banana contact is boxed.
READING 172 · 6 ORIGINAL VISUALS

AGRA

AGRA aligns the video features read by a separate action decoder with frozen semantic targets, improving manipulation while retaining multi-depth predictive guidance.

Read illustrated report →
Architecture diagram: downsampled future frames and a detailed first frame enter a VAE; a 30-layer WAN teacher is sliced into a 12-layer compact expert; the inference stack reuses video KV caches while actions continue denoising.
READING 173 · 6 ORIGINAL VISUALS

Efficient-WAM

A compact video expert can guide robot actions with coarse futures and sparse updates, but the fastest configuration sacrifices some task success and requires careful latency accounting.

Read illustrated report →
A 30-block main video model sends features from layers 4, 12, 20, and 30 through an MLP to three chained MCP modules, each with a noisy future input and flow-matching loss.
READING 174 · 6 ORIGINAL VISUALS

Next Forcing

Multi-chunk supervision improves a causal video/action backbone, while optional two-chunk generation trades deployment accuracy against a claimed reduction in video-denoising cost.

Read illustrated report →
Three stage panels show video-to-flow tokenization, supervised latent action prediction, and a state representation connected to write gate, memory bank, read gate and action head.
READING 175 · 6 ORIGINAL VISUALS

HiMem-WAM

Flow-supervised motion latents and discovered skill boundaries support a causal, memory-augmented robot policy, with stronger evidence for Stage II pretraining than for the isolated value of memory gating.

Read illustrated report →
Two transformer branches connected by observation-guided K/V routing, with rolling video memory and separate training and inference attention masks.
READING 176 · 6 ORIGINAL VISUALS

AHA-WAM

AHA-WAM reuses a slow video planner's latent context for fast, observation-conditioned action updates, gaining speed through both its asynchronous interface and substantial deployment optimization.

Read illustrated report →
Three-stage overview: egocentric video trains a Video DiT; its weights are copied into a video-motion pair with embodiment tags; retargeted teleoperation supplies whole-body fine-tuning data and motion/end-effector outputs.
READING 177 · 6 ORIGINAL VISUALS

MotionWAM

MotionWAM turns a single pass through a video world model into whole-body action features, improving G1 task success while leaving its precise motion-token encoding unresolved.

Read illustrated report →
A green full-computation path runs L DiT blocks and writes their residual to a cache. An orange cached path bypasses the blocks and adds the stored residual to h_0. A lower schedule aligns individual denoising steps between full and reused chunks.
READING 178 · 5 ORIGINAL VISUALS

C³ache

C³ache reuses an action transformer's residual across control cycles, reducing inference cost when the cache remains useful but risking large success losses when it becomes stale or replaces late denoising steps.

Read illustrated report →
A shared Dream-Tac block receives clean and noisy tactile, image, state and action tokens. Insets show DiT blocks, a tactile-change gate and factorized contact-aware attention.
READING 179 · 6 ORIGINAL VISUALS

Dream-Tac

Dream-Tac jointly denoises actions and visuo-tactile futures, using a deterministic contact-change bias to improve real robot success at the cost of an enlarged diffusion model.

Read illustrated report →
Language and current observation enter a video world model; generated frames feed an action extractor, then kinodynamic planning and UAV execution.
READING 180 · 6 ORIGINAL VISUALS

ImagineUAV

ImagineUAV converts an imagined visual route into flight through separate motion extraction and kinodynamic optimization, improving navigation success while retaining substantial generation latency.

Read illustrated report →
Three-column system pipeline: scripted and PICO demonstrations with whole-body control; Isaac Sim replay and success-only filtering with domain randomization; three evaluation levels.
READING 181 · 5 ORIGINAL VISUALS

SIMPLE

SIMPLE separates humanoid contact physics from photorealistic rendering to make demonstrations and policy comparisons reusable, at the cost of expensive rendering and still-limited evidence for broad physical transfer.

Read illustrated report →
Black shared-backbone arrows connect transformer blocks and a residual adapter; blue arrows lead to latent video, while red current-observation and pooled-state arrows lead to the StateFusionActionExpert.
READING 182 · 6 ORIGINAL VISUALS

Light-WAM

Light-WAM retains future-video supervision during training and directly decodes actions from pooled backbone states, gaining efficiency while giving up success on harder tasks.

Read illustrated report →
An episode strip marks fine manipulation in blue and subtask transitions in red. Below, motion parsing supplies manipulation labels and candidate windows; Qwen verification and boundary fusion supply per-frame subtask labels.
READING 183 · 6 ORIGINAL VISUALS

AdaWAM

AdaWAM gates subtask-language updates and visual foresight independently to improve difficult manipulation, trading modest per-step overhead against more successful task execution.

Read illustrated report →
An autoregressive Transformer takes encoded historical/current images, memory, instruction and meta-queries; physical-dynamics features feed separate Action and World Experts. A second panel scores multiple imagined futures.
READING 184 · 6 ORIGINAL VISUALS

WLA-0

WLA-0 trains language-guided dynamics features with a World Expert, allowing efficient action generation without future-image synthesis while leaving optional imagined-future selection and broad video transfer less fully established.

Read illustrated report →
Three panels show different video/action noise distributions, student and EMA-target consistency losses, and a shared-backbone student generating video and actions.
READING 185 · 6 ORIGINAL VISUALS

Flash-WAM

Matching consistency functions to video and action noise regimes makes a shared world-action model much faster, with lower task success and unresolved baseline-reporting conflicts.

Read illustrated report →
Training diagram with image/video VAE, action and state encoders, noised video/action tokens, a shared DiT, anchor-map controller, frozen text encoder and teacher, Gaze supervision, and video/action decoders.
READING 186 · 6 ORIGINAL VISUALS

Donk

Donk reuses one video-action denoising core for observed-scene prediction and text-conditioned paired data generation, with initial hand-camera geometry and an asymmetric visual-to-action information path.

Read illustrated report →
Video and action DiT branches share conditioning; future RGB, depth and semantic outputs sit above a dashed deployment region containing the initial frame, instruction, noisy actions and action decoder.
READING 187 · 6 ORIGINAL VISUALS

GeoSem-WAM

Training a video/action model to predict future geometry and semantics improves reported manipulation success while allowing those dense prediction heads to be removed at deployment.

Read illustrated report →
Three language-described manipulation events with unequal durations above a global-instruction history window; lower diagrams compare event-localized and fixed-length prediction targets.
READING 188 · 6 ORIGINAL VISUALS

WALL-WM

Event-aligned pretraining links language and visual futures to executable trajectories, improving reported instruction generalization while retaining expensive annotation and multi-module training.

Read illustrated report →
Diagram of visual and language inputs, foundation action models, a trajectory velocity field, and relative weights highlighting selected manipulation steps.
READING 189 · 6 ORIGINAL VISUALS

AttenA+

AttenA+ gives slow demonstrated actions greater relative training weight, improving several manipulation policies while depending on a task-specific link between speed and precision.

Read illustrated report →
Three stacked WAM schematics: a shared model emits actions and futures, an imagination module precedes an action decoder, and an auxiliary future branch is crossed out at inference. A legend distinguishes solid, dashed and dotted paths.
READING 190 · 6 ORIGINAL VISUALS

Beyond Task Success

Rollout diagnostics and sparse feature analysis expose architecture-dependent WAM behavior and cost, while leaving the causal value of future prediction unresolved.

Read illustrated report →
The architecture routes encoded observations, ego actions, ego state, and VLM text into one video diffusion transformer, with separate video/action history and a memory-selection stage.
READING 191 · 6 ORIGINAL VISUALS

DriveWAM

DriveWAM predicts ego motion through a generated future in one video-action transformer, gaining benchmark accuracy while retaining dependence on semantic guidance and the quality of its imagined future.

Read illustrated report →
Three connected modules: Qwen and DA3 feature encoding, a pose-supervised trajectory transformer, and a state-conditioned action transformer; demonstration arrows supply separate trajectory and action losses.
READING 192 · 6 ORIGINAL VISUALS

OASIS

Supervising future end-effector poses makes decoder features useful for rigid-body control, but successful execution still depends on a learned action decoder and feedback.

Read illustrated report →
Three panels show sliding-window track targets, image/action/track branches meeting a unified DiT, and 3D patchification of the query-point grid.
READING 193 · 6 ORIGINAL VISUALS

JOPAT

JOPAT couples future pixels, point tracks and robot actions in one denoiser, improving motion-sensitive manipulation while retaining the limits of sparse 2D supervision.

Read illustrated report →
Two-panel architecture: instruction-derived memory conditions a VLM with a VAE visual branch and a KV-conditioned action expert; the enlarged module shows hashing, concatenation, gating, convolution and pre-attention residual injection.
READING 194 · 6 ORIGINAL VISUALS

Key-Gram

Key-Gram gives a VLA policy addressable linguistic memory that improves compositional manipulation, while leaving parser reliability and the mechanics of protected knowledge expansion unresolved.

Read illustrated report →
Two policy diagrams share four colored annotation aspects. The VLA produces continuous actions through an action head. The WAM passes observations through video and joint encoders into DiT blocks, with separate video-decoder and post-training action-head branches.
READING 195 · 6 ORIGINAL VISUALS

DeMiAn

Re-annotating fixed demonstrations with task-appropriate language improves simulated policy learning, but deployment depends on how captions are selected, generated and injected.

Read illustrated report →
Time-unrolled diagram: observed image segments and a navigation instruction enter a latent autoregressive backbone; predicted latent blocks feed action decoders, with executed actions and observations connected through the environment.
READING 196 · 6 ORIGINAL VISUALS

WorldVLN

WorldVLN decodes predicted latent futures into drone waypoints and refreshes its history after execution, improving short-range navigation while retaining a large server-side backbone.

Read illustrated report →
Three bottom-to-top diagrams show VLA mapping current observation and language to action, WAM mapping them to next observation and action, and WM mapping current observation and action to next observation. Overlapping circles compare WAM, VAM, and Video Policy.
READING 197 · 6 ORIGINAL VISUALS

World Action Models survey

The survey organizes predictive robot policies by how future states and actions are coupled, while exposing the lack of controlled evidence that distinguishes useful foresight from auxiliary supervision.

Read illustrated report →
Training diagram with student/teacher vision encoders, resamplers, a World Predictor, Action Denoiser and Action Head; lower strip alternates prediction and denoising with selectable observation input.
READING 198 · 6 ORIGINAL VISUALS

DAWN

DAWN repeatedly couples action-conditioned latent prediction with world-conditioned trajectory denoising, gaining planning quality at an inference cost that depends on rollout horizon and interaction count.

Read illustrated report →
Two manipulation sequences and a table comparing transit and interaction success for Imagine-then-Execute and Joint Modeling under ID and three OOD settings.
READING 199 · 6 ORIGINAL VISUALS

HarmoWAM

HarmoWAM uses a learned stage gate to combine video-guided transit with latent-conditioned action diffusion for precise interaction, gaining robustness to controlled distribution shifts at the cost of video generation and a multi-module training pipeline.

Read illustrated report →
Motus action, video, and understanding experts feed an FFDC Transformer; a timeline aligns real observations, predicted actions and video, alongside a modality attention mask.
READING 200 · 5 ORIGINAL VISUALS

FFDC-WAM

FFDC uses a WAM’s imagined visual future to decide when to interrupt action execution, trading additional verification and corrective inference for better task success.

Read illustrated report →
Upper panel alternates action inference and execution along time. Lower panel shows three per-round projections of sampled chunks, teal successes, orange failures, selected stars and an inset of three-dimensional end-effector paths.
READING 201 · 6 ORIGINAL VISUALS

KeyStone

KeyStone selects a representative action chunk from parallel samples without retraining, gaining simulated task success when useful behaviors dominate and the GPU has spare sampling capacity.

Read illustrated report →
Three-part architecture diagram: joint self-attention between video and action experts in Stage 1, frozen experts and simulator-reward feedback in Stage 2, and the GPN encoder, fusion, Transformer, and squashed-Gaussian head.
READING 202 · 6 ORIGINAL VISUALS

NoiseGate

NoiseGate improves joint video–action control by learning when to denoise each imagined frame, at the cost of simulator rollouts for reward-based scheduler training.

Read illustrated report →
Three branch-selection diagrams compare predicted values, execution-based consistency and consensus agreement; each example retains branch 2. The legend distinguishes observed states, predicted actions, predicted states and values.
READING 203 · 6 ORIGINAL VISUALS

Dynamic Consistency in WAMs

Agreement among imagined futures can improve action selection without a value model, but predictable or widely shared futures can still represent failure.

Read illustrated report →
Action and current-state tokens produce fixed conditions for a latent flow network; a frozen residual encoder provides velocity targets and a frozen decoder produces future features.
READING 204 · 6 ORIGINAL VISUALS

RLA-WM

Compressing visual change makes generative feature dynamics economical, but stronger prediction metrics do not guarantee better policies on every robot.

Read illustrated report →
Multimodal inputs become image, object-slot, text, state and past-action tokens in a shared backbone, which branches into image-VQ, world and action heads.
READING 205 · 6 ORIGINAL VISUALS

OA-WAM

Persistent object addresses help the shared world/action model withstand geometric scene shifts, but the benefit depends on perception and does not establish universally robust control.

Read illustrated report →
Architecture diagram showing frozen teacher feature extraction, compression, a trainable router and adapters, context aggregation, and injection into a frozen student's conditioning stream.
READING 206 · 6 ORIGINAL VISUALS

CKT-WAM

A learned context interface improves a frozen student's manipulation performance, but keeps the teacher in the inference loop and leaves important adapter implementation details unresolved.

Read illustrated report →
Three-part architecture: robot states and camera calibration become KVAF images; a shared VAE encodes RGB, frame differences and KVAFs; video and KVAF DiT branches exchange information through an event-prediction and gate module.
READING 207 · 6 ORIGINAL VISUALS

EA-WM

Camera-aligned kinematic fields and event-supervised fusion improve simulated robot-video rollouts, but the visual action representation remains difficult to invert into precise numerical control.

Read illustrated report →
Three bottom-to-top diagrams compare a video generator followed by inverse dynamics, a shared observation/action backbone, and specialized video/action experts joined by attention. A legend identifies observations, actions and token colors.
READING 208 · 6 ORIGINAL VISUALS

World Models for Robot Learning

The survey maps how prediction becomes useful for robot action, while showing why architectural integration and visually plausible futures do not alone establish reliable control.

Read illustrated report →
Packed Mixture-of-Transformers architecture with shared context, prior latent queries, posterior future embeddings, two flow-matching action paths, alignment and regularization losses, and an isolated dual-branch attention mask.
READING 209 · 5 ORIGINAL VISUALS

Being-H0.7

Being-H0.7 uses future-informed latent alignment to train a direct robot policy, gaining strong whole-system performance while leaving the contribution of each training ingredient unisolated.

Read illustrated report →
Three parallel language, video and action streams with encoders, modality blocks and a shared middle multimodal-attention bridge; video and action have separate timestep inputs.
READING 210 · 6 ORIGINAL VISUALS

Motubrain

Motubrain learns aligned video and robot actions, then limits repeated visual computation to make chunked control practical, although its autoregressive mask specification is inconsistent.

Read illustrated report →
Left: RGB, state and action tokens enter shared DiT blocks, then separate main and depth branches. Right: noise sampling regions and action/video denoising trajectories.
READING 211 · 6 ORIGINAL VISUALS

X-WAM

X-WAM couples geometric supervision with early action decoding, improving the reported quality/latency tradeoff while retaining the delays and context limits of a large video-based controller.

Read illustrated report →
Two attention matrices compare current-only and future-visible action queries; training uses a shared backbone and residual adapter, while inference discards the teacher. The original training panel contains conflicting stop-gradient and loss markings.
READING 212 · 6 ORIGINAL VISUALS

PFD

PFD distills the action-velocity change induced by privileged future observations into a current-only policy, improving manipulation success while retaining a small, measurable adapter cost.

Read illustrated report →
Diagram of the pi_0.6 VLA, its action expert, and an encoder–decoder branch compressing image embeddings into one 2048-dimensional RL token.
READING 213 · 6 ORIGINAL VISUALS

RL Token

RLT turns a frozen VLA’s features and action proposals into a lightweight online refinement interface, improving precision-task execution while retaining human supervision and a strong behavioral anchor.

Read illustrated report →
Scene reconstruction feeds a simulator, while predicted human and robot videos follow separate paths through hand retargeting or inverse dynamics to embodied validation.
READING 214 · 6 ORIGINAL VISUALS

RoboWM-Bench

RoboWM-Bench turns generated manipulation videos into executable robot trajectories, gaining an operational test of physical feasibility while inheriting errors from action extraction and simulation.

Read illustrated report →
Architecture diagram: a human or high-level policy supplies subtask language to the VLA and BAGEL world model; generated goal images join observation memory and metadata before the action expert produces commands.
READING 215 · 6 ORIGINAL VISUALS

π0.7

Detailed context lets π0.7 learn from heterogeneous robot experience and accept generated visual goals, improving steerability while adding annotation, runtime-compute and evaluation burdens.

Read illustrated report →
CLWM diagram showing frozen DINOv3 and language encoders, historical latent features and actions, video and action model blocks, and shared TTT memory.
READING 216 · 6 ORIGINAL VISUALS

DexWorldModel

CLWM uses predicted semantic features to guide actions and speculative computation, while its strong manipulation results leave memory fidelity, latency–accuracy tradeoffs, and several implementation details unresolved.

Read illustrated report →
Two world-action model schematics. AIM adds Future Value to Future Video and Action, with an arrow labelled Self-Distillation from Future Value to Action.
READING 217 · 5 ORIGINAL VISUALS

AIM

AIM uses predicted contact maps to connect visual foresight to robot actions, trading explicit simulator supervision for an interpretable spatial interface whose causal benefit remains incompletely tested.

Read illustrated report →
Two diagrams show inference above and training below: video noise enters a diffusion transformer; its clean latent goes to a decoder and through 3D pooling to an action U-Net. Training marks the encoder frozen and the video-to-action connection detached.
READING 218 · 6 ORIGINAL VISUALS

VAG

VAG uses synchronized video and action denoising to synthesize robot-training pairs, with useful replay and transfer results but no isolated test of its global-pooling mechanism.

Read illustrated report →
Four panels show visual-token pretraining, multi-task supervised fine-tuning, reinforcement-learning policy updates and an inference response ordered from perception through visual tokens to action and waypoints.
READING 219 · 5 ORIGINAL VISUALS

VLA-World

VLA-World conditions future-image tokens on a short-term motion prediction and uses them to refine a driving plan, improving offline nuScenes metrics while leaving the accuracy and causal usefulness of imagined evidence unresolved.

Read illustrated report →
Architecture diagram showing navigation, camera, action, world-query and action-query inputs to an LLM, separate action and video outputs, and an optional feedback loop.
READING 220 · 6 ORIGINAL VISUALS

LMGenDrive

LMGenDrive uses joint action and future-video training to improve a language-conditioned driving policy, then removes video generation from online control while retaining it for offline rollouts.

Read illustrated report →
Human collection and robot deployment branches converge on geometry processing and a point-cloud policy, which sends object-motion targets to a manipulation arm and camera-motion targets to a perception arm. A return arrow connects perception movement to new observations.
READING 221 · 6 ORIGINAL VISUALS

ActiveGlasses

ActiveGlasses learns object and camera motion from bare-hand demonstrations, trading direct hand retargeting for calibrated 3D perception and task-specific trajectory learning.

Read illustrated report →
STAR pipeline with reference/history cache above a DiT block, user-camera and depth-warping inputs at left, decoded video at right, and motion/perceptual distillation teachers below.
READING 222 · 6 ORIGINAL VISUALS

INSPATIO-WORLD

Reference-anchored autoregressive diffusion and geometric reprojection support controllable video roaming, while dual-teacher distillation targets visual fidelity and generated-region memory remains incomplete.

Read illustrated report →
A robot gripper projects three semantic points into red position, green normal and blue up-point heatmap channels; the blue channel also stores gripper openness.
READING 223 · 6 ORIGINAL VISUALS

Action Images

Encoding robot pose and gripper state as multiview images lets one video backbone generate actions, while geometric decoding and open-loop execution remain critical limits.

Read illustrated report →
Pipeline from an observation and candidate instructions through a WAM, action-trajectory rendering, risk classification, and selective simulator execution with a final human judge.
READING 224 · 6 ORIGINAL VISUALS

JailWAM

JailWAM turns predicted robot actions into visual risk evidence, reducing simulator use at the cost of missing some hazards that emerge during closed-loop execution.

Read illustrated report →
History images and optional text are encoded on the left; a unified diffusion transformer produces future video and trajectory above two noisy target groups. A lower strip shows successive video continuation chunks.
READING 225 · 6 ORIGINAL VISUALS

DriveVA

DriveVA couples future video and trajectory generation in one transformer, improving reported planning and transfer while leaving the model vulnerable to mutually consistent but incorrect futures.

Read illustrated report →
Three-panel architecture showing virtual projection of a point cloud and end-effector pose, a shared multiview video diffusion transformer, heatmap back-projection, and a separate rotation/gripper predictor with global and local latent features.
READING 226 · 6 ORIGINAL VISUALS

SpatialVAM

SpatialVAM makes robot positions look like multiview video, gaining strong low-demonstration control through shared RGB/heatmap prediction while retaining separate action decoding and substantial inference cost.

Read illustrated report →
Three parallel understanding, perception and action pathways connect separate QKV and FFN blocks through masked joint attention. Inputs include a prompt, camera views, sparse anchors and noised actions; outputs include language, perception and trajectories.
READING 227 · 6 ORIGINAL VISUALS

UniDriveVLA

UniDriveVLA lets three specialized experts exchange information inside one model, improving selected driving and understanding metrics while leaving comfort, motion forecasting and full semantic retention unresolved.

Read illustrated report →
Language, multi-view images, current action and learned queries enter a shared LLM; world and action embeddings condition separate depth, video and trajectory transformers.
READING 228 · 6 ORIGINAL VISUALS

DriveDreamer-Policy

Depth-supervised queries enrich a shared driving representation for video and trajectory prediction, while separate generators let planning proceed without rendering an imagined world.

Read illustrated report →
Three time steps show observation encoders, posterior and prior latent states, recurrent transitions, observation decoders, reward estimators and a diffusion-policy fine-tuning pathway. Adjacent encoder arrows meet at predicted actions.
READING 229 · 6 ORIGINAL VISUALS

WAM

Inverse-action supervision can improve the frozen features used by a diffusion policy, but this paper's inconsistent result summaries and absent ablations leave the size and cause of that benefit uncertain.

Read illustrated report →
Two annotated photographs show the tabletop bimanual setup with master and follower arms, cameras and lighting, and the mobile robot with cameras, lifting column and omnidirectional base.
READING 230 · 6 ORIGINAL VISUALS

ManipArena

Controlled physical trials expose how demonstration selection, language annotations and pretraining provenance change manipulation rankings, while limited coverage and inconsistent reporting constrain the conclusions.

Read illustrated report →
Two-panel architecture showing history, ego and prompts converted to depth-fused embeddings, followed by alternating future-frame outputs and action queries through Uni-World VLA.
READING 231 · 6 ORIGINAL VISUALS

Uni-World VLA

Uni-World VLA alternates imagined frames and waypoint predictions in a shared backbone, improving reported NAVSIM planning while leaving the causal contribution of feedback and deployment behavior unresolved.

Read illustrated report →
Architecture with blue vision tokens, cyan text tokens with yellow outlines, orange action tokens, green hatched image-generation tokens, separate understanding/action/generation modules and a shared causal-attention block.
READING 232 · 6 ORIGINAL VISUALS

Vega

Vega couples instruction-conditioned trajectory denoising to future-image prediction, gaining dense training supervision while leaving general instruction-following reliability unmeasured.

Read illustrated report →
Two-panel diagram: a separate robot policy supplies actions to an autoregressive world model above; shared-prefix sampling, candidate rewards and positive/negative denoising branches form the post-training loop below.
READING 233 · 6 ORIGINAL VISUALS

PersistWorld

Training on ranked continuations of self-generated histories improves robot-video persistence, but increases sampling cost and leaves physical validity dependent on the reward.

Read illustrated report →
Three-stage training diagram: historical images and actions feed a trajectory policy and vocabulary sampler; candidate-specific latent rollouts produce temporal rewards that feed group advantages and policy optimization.
READING 234 · 6 ORIGINAL VISUALS

DreamerAD

DreamerAD trains a driving policy with fast latent rollouts and learned prefix rewards, gaining NAVSIM planning score while constraining exploration and accepting lower ego progress.

Read illustrated report →
Architecture with a compressive image encoder and geometric alignment above, an EMA-supervised dynamic predictor below, and a separate command-conditioned trajectory decoder at right.
READING 235 · 6 ORIGINAL VISUALS

Latent-WAM

Geometry distillation and causal future-state prediction train compact driving representations, while deployment uses a direct trajectory decoder with no world-model rollout.

Read illustrated report →
Real and simulated driving data sit above a pipeline that converts 2D paths through spatial, agent and map attention into 6-DoF paths, then projected layouts and multi-view video.
READING 236 · 5 ORIGINAL VISUALS

PhyGenesis

Learning to repair impossible trajectories and learning to render difficult physical interactions are complementary, but their measured gains depend on synthetic supervision and a stylized evaluation domain.

Read illustrated report →
Three groups of camera and tactile VAE latents enter a video model; an arrow carries their representation into a separate action diffusion branch, with both branches marked ×28.
READING 237 · 5 ORIGINAL VISUALS

VTAM

VTAM makes tactile deformation part of predictive video modeling and action supervision, with strong task success but unresolved implementation and mechanism claims.

Read illustrated report →
Eight images of a robot and service bell show the clean scene and camera, robot-state, language, light, background, noise and layout perturbations, with sub-dimension labels.
READING 238 · 6 ORIGINAL VISUALS

WAM Robustness Study

Video-based policies show strong visual robustness, but benchmark dependence, unresolved score aggregation and costly inference prevent a universal WAM advantage.

Read illustrated report →
Seven camera inputs pass through an encoder, noised video tokens, view-temporal attention, condition injection, and a decoder to seven camera outputs.
READING 239 · 6 ORIGINAL VISUALS

X-World

X-World adapts a controllable multi-camera video generator into a streaming driving simulator, but its qualitative demonstrations leave physical fidelity and policy-evaluation reliability unmeasured.

Read illustrated report →
Observation and instruction encoders feed a video predictor with one stochastic transition and deterministic completion; predicted and expert latents enter L1 and cosine reward blocks.
READING 240 · 6 ORIGINAL VISUALS

VAMPO

Rewarding completed video latents can improve the early predictive features used by a separate action policy, but most measured control gains require retraining that policy.

Read illustrated report →
Two observation-and-instruction pipelines pass generated videos through an inverse dynamics model. The lower pipeline adds an IDM-derived reward arrow returning to the video generator.
READING 241 · 6 ORIGINAL VISUALS

EVA

EVA improves a video planner by rewarding smooth, limit-respecting IDM-decoded actions, but this kinematic proxy can favor motion that never completes the task.

Read illustrated report →
Two panels show joint video/action training through causal DiT blocks and inference with an optional future-video stream and KV cache.
READING 242 · 6 ORIGINAL VISUALS

GigaWorld-Policy

A shared causal Transformer uses future-video supervision to train a robot policy whose deployed actions do not require future-video generation.

Read illustrated report →
Three-stage pipeline: a VLM proposes robot actions, collected trajectories train a separate video world model, and positive/negative predicted outcomes feed an ORPO update to the VLM.
READING 243 · 6 ORIGINAL VISUALS

DreamPlan

DreamPlan moves expensive video-based action verification into offline preference training, improving a deformable-manipulation planner's reported scores while allowing direct action inference at deployment.

Read illustrated report →
Three columns compare joint video/action denoising, video-then-action generation, and Fast-WAM, with separate training and inference rows and KV-cache arrows.
READING 244 · 6 ORIGINAL VISUALS

Fast-WAM

Future-video supervision can improve an action policy even when deployment uses only current-observation features, avoiding video denoising while retaining iterative action sampling.

Read illustrated report →
Architecture linking a reference image and CLIP instruction embedding to video diffusion features, geometric and semantic decouplers, a Uni-Perceiver and a diffusion action policy.
READING 245 · 6 ORIGINAL VISUALS

S-VAM

S-VAM distills generated-video geometry and semantics into first-step foresight that improves manipulation while adding a modest latency cost over raw-feature control.

Read illustrated report →
Two stacked pipelines show video diffusion above and planning below. Frozen encoders recur in the lower pipeline, and the world model inside the rewarder is explicitly marked train-only.
READING 246 · 6 ORIGINAL VISUALS

WorldDrive

WorldDrive transfers trajectory-conditioned video representations into a separate planner and distills future latents into a fast candidate scorer, trading online video simulation for learned future estimation.

Read illustrated report →
A two-branch JEPA diagram with image/video patchifiers, a masked context encoder, multi-level fusion and predictor above, and an unmasked EMA target encoder below. Blue context and orange masked predictions meet separate L1 losses through stopped-gradient target paths.
READING 247 · 6 ORIGINAL VISUALS

V-JEPA 2.1

Supervising visible patches at multiple encoder depths makes video features substantially more useful for dense vision, while extra supervision and scaling recover much of the recognition performance lost to local reconstruction.

Read illustrated report →
Two frames pass through an encoder. The current embedding and action feed a predictor; next-embedding MSE and two SIGReg branches train the model. An inset illustrates random projections and Gaussian distribution matching.
READING 248 · 6 ORIGINAL VISUALS

LeWorldModel

LeWM trades pixel reconstruction and elaborate anti-collapse losses for Gaussian-regularized latent prediction, enabling compact planning with task-dependent control and representation limits.

Read illustrated report →
Three proxy-training schematics and two success-rate plots compare grounding, FLARE-style feature alignment and video generation; orange video curves exceed the green and blue curves.
READING 249 · 6 ORIGINAL VISUALS

DiT4DiT

Video-generation features can condition effective robot actions after one backbone pass, but the resulting dual-transformer policy trades deployment speed for task performance.

Read illustrated report →
Two training stages connect frozen world-model latents and VLA actions through adapters. A right-hand inset shows the residual Transformer and the addition of its output to the base action latent.
READING 250 · 6 ORIGINAL VISUALS

World2Act

A frozen video world model can teach a small residual controller through aligned action latents, improving manipulation without pixel-derived action labels while retaining a difficult timing and contact gap.

Read illustrated report →
Two functional experts share conditional attention: camera BEV, language and soft world prompts feed the world-language branch; ego state and noisy actions feed planning. Rendering is marked pretraining-only.
READING 251 · 6 ORIGINAL VISUALS

SAMoE-VLA

SAMoE-VLA uses scene-conditioned parameter fusion to improve a world-informed driving planner, with accuracy gains that require separate scrutiny of collision behavior and implementation consistency.

Read illustrated report →
Two policy diagrams joined by a subtask instruction: the left updates text memory, while the right encodes video history and produces continuous actions through an action expert.
READING 252 · 6 ORIGINAL VISUALS

MEM

MEM combines compressed semantic history with dense recent video so a robot can remember task progress and adapt its manipulation, at the cost of summary supervision and an incompletely specified training recipe.

Read illustrated report →
Front-view, instruction and navigation branches enter a multimodal language model; endpoint prediction is followed by interpolation and parallel trajectory refinement.
READING 253 · 6 ORIGINAL VISUALS

LinkVLA

LinkVLA couples language and driving trajectories through one token decoder, then accelerates action generation with endpoint-conditioned refinement whose reported latency excludes textual reasoning.

Read illustrated report →
Two-stage diagram: vision, language and proprioception feed a shared DiT with intermediate progress/state heads and a final action head; a frozen base then supplies actions and predictions to residual reinforcement learning.
READING 254 · 6 ORIGINAL VISUALS

SC-VLA

Shared progress and end-effector predictions improve a flow policy and guide a separate residual controller, but the demonstrated online gains rely on simulation interaction and sparse environment reward.

Read illustrated report →
Matched head and wrist cameras for human capture beside a diagram connecting text and visual encoders, a pretrained VLM, noisy actions, a DiT expert and an action decoder.
READING 255 · 6 ORIGINAL VISUALS

EgoScale

Explicit human wrist and hand supervision improves dexterous robot learning at scale, while matched human–robot experience remains essential for efficient adaptation.

Read illustrated report →
System diagram with two depth cameras, depth-to-normal conversion, robot state, a demonstration-seeded replay buffer, random image augmentation, a world model and a separate RL policy feeding robot actions.
READING 256 · 5 ORIGINAL VISUALS

Learning to Unfold Cloth

Geometry-focused observations and demonstration-seeded, augmented replay make a DreamerV2 cloth controller more effective, while depth-sensor limits and garment-specific training constrain transfer.

Read illustrated report →
Two training stages: an encoder–FSQ–decoder tokenizer, followed by current/next-frame factorizers, per-slot inverse and forward dynamics, and an aggregator with a current-feature bypass.
READING 257 · 6 ORIGINAL VISUALS

FLAM

FLAM gives each learned scene factor its own latent action while sharing interaction-aware dynamics, improving multi-entity video modeling but still requiring action supervision to train an executable policy.

Read illustrated report →
Video diffusion diagram with reference and memory frames, noisy future frames, two action-conditioning paths, a decoder and an autoregressive feedback arrow.
READING 258 · 6 ORIGINAL VISUALS

WoVR

WoVR makes imagined reinforcement learning more useful by stabilizing the simulator, shortening unreliable prefixes and refreshing its policy coverage, while leaving residual model error and a distinct physical-robot data protocol.

Read illustrated report →
Architecture diagram with four task-conditioned distributions, action and DINO streams, modality register tokens, shared self-attention, and separate cross-attention and feed-forward branches.
READING 259 · 6 ORIGINAL VISUALS

LDA-1B

LDA-1B turns mixed-quality embodied recordings into policy and latent-dynamics supervision through one task-conditioned generator, gaining data efficiency while leaving deployment details and some causal attributions unresolved.

Read illustrated report →
Two separate model blocks: the lower world model supplies video-state tokens and values, the upper VLA emits robot actions, and human-in-the-loop rollout data feed continual training.
READING 260 · 6 ORIGINAL VISUALS

GigaBrain-0.5M*

RAMP trains a VLA with predicted future latents and value-derived improvement labels, gaining task success while depending on corrected robot experience and incompletely specified evaluation protocols.

Read illustrated report →
Two training branches initialize a blue Genie Envisioner dynamics model and a green π0.5 value model; multiview predictions flow between them and provide feedback for policy optimization.
READING 261 · 6 ORIGINAL VISUALS

RISE

RISE trains a robot policy on action-conditioned imagined futures scored by a separate value model, exchanging physical exploration for compute while retaining a substantial real-data anchor.

Read illustrated report →
Architecture diagram showing prefix-masked query compression, quantized latent codes, a causal decoder, two stop-gradient markers and direct and ControlNet paths into a pretrained video diffusion model.
READING 262 · 6 ORIGINAL VISUALS

VideoWorld 2

A diffusion appearance prior helps compact dynamics codes transfer across visual settings, while autoregressive code prediction still requires separate supervision to become a robot policy.

Read illustrated report →
Two stacked training diagrams show VLM latent actions conditioning a world predictor, frozen video-state targets for alignment, and an additional robot action head in the lower panel.
READING 263 · 6 ORIGINAL VISUALS

VLA-JEPA

Predicting frozen video features helps a VLA learn representations that improve perturbation robustness, but the reported human-video benefit depends strongly on the evaluation setting.

Read illustrated report →
Three branches show world-model and inverse-dynamics data generation for VLA training, observation-and-instruction-conditioned action planning into a simulator, and repeated policy–world-model interaction followed by a VLM success judge.
READING 264 · 6 ORIGINAL VISUALS

WorldArena

WorldArena separates the quality of imagined robot videos from their usefulness for training, evaluating and executing policies, revealing that one video score cannot establish functional competence.

Read illustrated report →
Four connected stages: success and near-success data collection, diffusion video and reward training, a world-model/VLA rollout loop, and physical deployment returning new trajectories to the dataset.
READING 265 · 6 ORIGINAL VISUALS

World-VLA-Loop

Learning subtle failures and a shared video–reward representation makes a video simulator more useful for VLA reinforcement learning, while policy-driven data refresh addresses new simulator exploits at the cost of further target-environment interaction.

Read illustrated report →
Exocentric images pass through a locked Cosmos tokenizer, four tactile images through Sparsh-X, and projected blue and purple tokens enter an action-conditioned transition predictor with separate visual and tactile outputs.
READING 266 · 5 ORIGINAL VISUALS

VT-WM

Predicting touch alongside vision improves contact-sensitive imagination and goal-image planning, but execution remains computationally expensive and evidence is confined to a narrow robot setting.

Read illustrated report →
XGen's eight numbered panels show human video, recovered human and humanoid motion, object-anchor configuration, contact refinement, physics-simulated non-contact trajectories and their assembled interaction sequence. Yellow markers identify augmentation operations.
READING 267 · 6 ORIGINAL VISUALS

HumanX

HumanX uses contact anchors and physics to expand human demonstrations into interaction references, then distills imitation policies for robot deployment, with generalization bounded by the synthesized data and available sensing.

Read illustrated report →
A language block and interleaved video/action blocks consume task text, observed frames and actions; generated video and action outputs appear above successive rollout stages.
READING 268 · 6 ORIGINAL VISUALS

LingBot-VA

LingBot-VA decodes robot actions from predicted visual futures inside a shared-attention model, trading costly video generation for partial denoising and feedback-grounded asynchronous execution.

Read illustrated report →
Three rows show raw multi-view images and blank placeholders becoming latent frames, injected proprioception/actions/value, and clean conditioning versus noisy prediction targets.
READING 269 · 6 ORIGINAL VISUALS

Cosmos Policy

A video diffusion model can learn executable actions through latent frame injection, while rollout-refined prediction improves action selection at a substantial inference cost.

Read illustrated report →
Multi-view images pass through a vision encoder and QT-Former; visual and instruction embeddings enter a LoRA-adapted LLM whose output layer serves trajectories, future images and VQA.
READING 270 · 6 ORIGINAL VISUALS

UniDrive-WM

Predicting an image after a driving plan supplies useful training feedback, while discrete and continuous visual decoders trade speed against image fidelity.

Read illustrated report →
Data construction diagram connecting in-house and public robot data to GPT selection, human labeling and benchmark examples, alongside five labeled distribution charts.
READING 271 · 6 ORIGINAL VISUALS

WoW-World-Eval

WoW-World-Eval connects generated robot videos to human judgment and IDM-mediated execution, exposing weak planning while making assessment depend on learned judges and calibration.

Read illustrated report →
CLAP diagram with visual inverse and forward dynamics, an action encoder and decoder, a contrastive similarity matrix, and NTP/RF policy variants.
READING 272 · 6 ORIGINAL VISUALS

CLAP

CLAP uses a robot-grounded token vocabulary to learn from human videos, then adapts a faster continuous controller while regularizing its semantic backbone.

Read illustrated report →
RGB-D scene observations and joint actions converted through a robot URDF form a concatenated point cloud; temporal robot features and DINOv3 scene features feed PTv3 to predict full-scene point flow.
READING 273 · 6 ORIGINAL VISUALS

PointWorld

PointWorld makes robot motion a geometric input to fast 3D scene prediction, enabling planner-based manipulation while depending on calibrated observations and accurately realized robot trajectories.

Read illustrated report →
Three Transformer branches for video generation, action and understanding share a central attention block; video and action timesteps feed their own AdaLN layers, and the upstream VLM has a freeze marker.
READING 274 · 6 ORIGINAL VISUALS

Motus

Motus unifies video and action generation through interacting experts and motion-based pretraining, with stronger average robot performance but unresolved attribution and reporting gaps.

Read illustrated report →
Two-stage training diagram: tactile, visual and language tokens enter a QwenLM decoder; sampled action responses are then marked chosen or rejected by comparison with ground truth.
READING 275 · 6 ORIGINAL VISUALS

VTLA

VTLA combines wrist vision and tactile histories in a token-generating policy, then improves OOD action accuracy through preference learning whose physical-control benefit remains unisolated.

Read illustrated report →
A green video transformer stack receives context frames, noisy future video and text cross-attention. A shallower blue action stack receives proprioception and noisy actions. The solid video-to-action arrow carries observed context; the dashed future-video arrow is labeled L_joint only.
READING 276 · 6 ORIGINAL VISUALS

Dyna-2

Human-video supervision supports cross-embodiment transfer while the scaling policy remains reactive; the reported gains depend on data scale, model variant and evaluation metric.

Read illustrated report →
Streaming camera, optional LiDAR and text inputs enter a shared Percept-WAM backbone; upper branches show PV perception, BEV perception, autoregressive language and parallel waypoint decoding, with a bidirectional KV-cache connection.
READING 277 · 6 ORIGINAL VISUALS

Percept-WAM

A shared VLM learns spatial perception tokens and trajectory queries, improving reported driving benchmarks while leaving important decoder, cache and training details in unsupplied appendices.

Read illustrated report →
Timeline of sparse past frames and noisy future frames entering spatial and temporal transformers, with three camera views and frame-aligned history/action pose embeddings.
READING 278 · 6 ORIGINAL VISUALS

Ctrl-World

Ctrl-World makes existing robot policies interact with generated multi-view observations, enabling targeted improvement from selected synthetic successes while retaining a substantial gap in precise physical execution.

Read illustrated report →
A four-row timeline places planner, world action model, simulator, and synthesizer systems across two yearly and four quarterly time bins.
READING 279 · 6 ORIGINAL VISUALS

World Models for VLA Agents

The survey classifies world models by how predicted futures enter an agent’s workflow, while its compiled scores leave causal gains and real-world reliability unresolved.

Read illustrated report →
Three task branches share the MonoDream backbone: action decoding on the left and current/future panoramic RGB-D feature supervision on the right. Green, orange and blended tokens mark text, vision and UNR.
READING 280 · 6 ORIGINAL VISUALS

MonoDream

MonoDream uses training-only panoramic feature prediction to improve a shared monocular navigation policy, trading richer supervision for a simple deployment input.

Read illustrated report →
Architecture with multi-view image backbone and BEV encoding, scene/agent/map query modules, separate visual projectors, textual ego and driver inputs, a driving VLA, and two command-conditioned planned paths. A right-hand legend identifies visual and textual tokens.
READING 281 · 6 ORIGINAL VISUALS

OpenDriveVLA

OpenDriveVLA teaches a shared autoregressive waypoint decoder through structured visual-language alignment and auxiliary agent forecasting, yielding competitive open-loop planning while leaving closed-loop safety untested.

Read illustrated report →
Green video-model blocks receive encoded frames and text; black arrows carry their features down into blue action-model blocks, which also receive action noise and command/ego-state context. Separate branches output future video and trajectory.
READING 282 · 6 ORIGINAL VISUALS

DriveLaW

DriveLaW conditions a separate diffusion planner on early video-model features, gaining planning performance while leaving the exact link between visual fidelity and action reliability unmeasured.

Read illustrated report →
Task language enters an object-library/API code-generation loop at left; expert trajectories and randomized scenes feed downstream policy training at right.
READING 283 · 6 ORIGINAL VISUALS

RoboTwin 2.0

RoboTwin 2.0 uses simulation-tested programs and randomized demonstrations to improve bimanual policy robustness, while broad task coverage and reliable visual error diagnosis remain unresolved.

Read illustrated report →
Left: colored low-level action tokens are compressed into stage tokens and repeated as aligned skills. Right: observation and instruction feed a high-level policy whose predicted skill conditions a low-level policy.
READING 284 · 4 ORIGINAL VISUALS

HiLAM

HiLAM learns variable-duration skills from pretrained motion latents to improve hierarchical policy adaptation, while retaining dependence on a pretrained extractor and action-labeled fine-tuning.

Read illustrated report →
Two consecutive human video frames enter a latent-action encoder. Its compact action code and the earlier frame enter a decoder that reconstructs the next frame. Retrieved human and robot frame pairs appear to the right.
READING 285 · 6 ORIGINAL VISUALS

DreamDojo (v1)

Continuous latent actions help transfer human-video experience into a robot simulator whose distilled student streams faster, while separate policy and value models are required to turn predictions into executed plans.

Read illustrated report →
Text, frame and action embeddings enter a shared causal Transformer. Repeated output groups feed frame and action diffusion decoders; an inset shows action denoising with cross-attention to physical-token conditioning.
READING 286 · 6 ORIGINAL VISUALS

PhysGen

PhysGen transfers video pretraining into a joint continuous video–action policy, with strong LIBERO results but unresolved control timing and limited causal evidence for physical understanding.

Read illustrated report →
An initial scene and task goal feed a repeated policy–world-model loop; generated observations return to the policy, and a separate VLM produces a reward. The left panel shows ordinary, edited-image and changed-instruction starts.
READING 287 · 6 ORIGINAL VISUALS

WorldGym

An action-conditioned video simulator can recover useful robot-policy rankings from initial images, but its scores inherit both dynamics and visual-grading errors.

Read illustrated report →
AutoMoT diagram with frozen 4B understanding expert, trainable 1.6B action expert, layer-wise joint attention, language/decision/planning heads, and a timeline reusing KV states between understanding updates.
READING 288 · 6 ORIGINAL VISUALS

AutoMoT

AutoMoT uses shared attention and a reusable scene cache to accelerate action prediction, trading fresh semantic context for faster replanning while keeping current action-side vision.

Read illustrated report →
Two-stage pipeline: a VLA policy collects real data to fine-tune a world model and generate synthetic data; real and synthetic trajectories then feed policy training with a reward model.
READING 289 · 6 ORIGINAL VISUALS

VLAW

VLAW turns real policy failures into better visual dynamics, then uses filtered imagined successes to improve a separate robot policy, with reliability limited by both model fidelity and reward selection.

Read illustrated report →
Three stages: a DINOv2 and NSVQ latent tokenizer with a forward reconstruction decoder; a shared video-language model predicting image and latent tokens; and noisy-action encoding with a flow decoder.
READING 290 · 6 ORIGINAL VISUALS

ViPRA

Joint future-image and motion-latent pretraining supplies useful robot-control priors, but deployment relies on action-only continuous decoding and gains depend on the benchmark.

Read illustrated report →
DrivingWorld pipeline from historical ego poses and front-view images, through separate tokenizers and two autoregressive modules, to predicted pose and image decoders.
READING 291 · 6 ORIGINAL VISUALS

DrivingWorld

DrivingWorld couples pose and video prediction through separate temporal and within-frame autoregression, improving reported generation while leaving important decoding and evaluation details unresolved.

Read illustrated report →
Three training stages show images, BEV, history and text entering InternVL3-2B, with denoiser, action and reward heads and trainable/frozen markers.
READING 292 · 6 ORIGINAL VISUALS

DriveWorld-VLA

Shared vision-language features support an action-conditioned BEV world model whose rewards refine a driving planner, with benchmark gains that still depend on staged training and incompletely specified inference details.

Read illustrated report →
Language, RGB images, end-effector pose and diffusion time condition a diffusion transformer; noisy action tokens enter below and denoised actions leave above, with an adaptive-normalization block enlarged at right.
READING 293 · 6 ORIGINAL VISUALS

TRI Large Behavior Models

Diverse action pretraining produces stronger task-specific manipulation policies, but its measured benefit depends on preprocessing, task-completion metrics and controlled evaluation.

Read illustrated report →
Two world-model branches above an interleaved language, image and action backbone: autoregressive visual tokens on the left and conditional latent denoising on the right. Snowflakes mark frozen modules and flames mark trainable modules.
READING 294 · 6 ORIGINAL VISUALS

DriveVLA-W0

Predicting visual futures during training improves driving representations at large data scale, while an optional action expert reduces decoding cost without generating images during control.

Read illustrated report →
Tracked upper-body skeleton and enlarged hand joints beside nine egocentric manipulation examples with colored fingertip trails.
READING 295 · 6 ORIGINAL VISUALS

EgoDex

EgoDex makes dexterous human motion a large-scale supervised learning target, but its strongest sampling results measure offline trajectory coverage rather than executable robot control.

Read illustrated report →
Diagram linking a front-view image, optional description and ego trajectory to a video generator, followed by direct video metrics and SLAM-based trajectory metrics.
READING 296 · 6 ORIGINAL VISUALS

DrivingGen

DrivingGen reveals tradeoffs between convincing driving videos and faithful ego motion, while making trajectory evaluation depend on imperfect motion reconstruction.

Read illustrated report →
Three-part WMPO loop: an old policy and world model alternate actions and frames, a reward model scores grouped trajectories, and a buffer supports a policy update.
READING 297 · 6 ORIGINAL VISUALS

WMPO

WMPO turns action-conditioned video generation into an on-policy training environment for a VLA, trading expensive environment rollouts for dependence on learned dynamics and reward accuracy.

Read illustrated report →
Two-panel diagram: adjacent-frame dynamics tokenization with separate ego/environment codebooks, and a causal policy supervised to generate dynamics before actions.
READING 298 · 6 ORIGINAL VISUALS

DynVLA

Predicting compact ego and environment dynamics before trajectory tokens improves reported driving benchmarks, but forecasting mistakes can become planning mistakes.

Read illustrated report →
Five image columns show a flower, geese, food, seals and sheep, followed by noisy pre-refinement and cleaner post-refinement patch cosine-similarity maps.
READING 299 · 6 ORIGINAL VISUALS

DINOv3

DINOv3 preserves spatial patch relationships during large-scale self-supervised training, producing versatile frozen features at the cost of substantial pretraining and task-dependent downstream adaptation.

Read illustrated report →
An image branches to a frozen foundation 3D model and an image tokenizer. Teacher features and intermediate VLA visual tokens feed an alignment block; image and text tokens continue through the VLA to action outputs.
READING 300 · 6 ORIGINAL VISUALS

Spatial Forcing

Spatial Forcing transfers geometric features into a VLA during fine-tuning, improving manipulation while keeping the base inference path, but leaving the geometry-to-action explanation only partly tested.

Read illustrated report →
Two frozen visual-encoder branches feed inverse dynamics; its latent action conditions a forward predictor, with three illustrated ways to limit latent information.
READING 301 · 6 ORIGINAL VISUALS

Latent Actions in the Wild

Regularized continuous latent actions turn natural-video prediction into a reusable planning interface, but that interface still needs visual context and supervised action mapping.

Read illustrated report →
Three autoregressive RWM-U columns produce observation predictions, uncertainty boxes, and penalized rewards for MOPO-PPO, with a separate policy block returning to the models.
READING 302 · 6 ORIGINAL VISUALS

RWM-U + MOPO-PPO

Ensemble disagreement makes long imagined rollouts useful for offline locomotion learning, but successful control still depends on penalty calibration and diverse data.

Read illustrated report →
Three-panel diagram: a unified model generates head and wrist observations and a reward answer; short imagined branches feed PPO training of a separate VLA; task photographs appear at right.
READING 303 · 6 ORIGINAL VISUALS

VLA-MBPO

VLA-MBPO improves a separate robot policy with short, multiview imagined rollouts, trading fewer environment interactions for world-model compute and residual prediction bias.

Read illustrated report →
Two-stage diagram showing RGB encoding and consistency decoding at left, autoregressive latent prediction in the center, and Conv3D, FiLM, spatial attention and temporal attention blocks at right.
READING 304 · 6 ORIGINAL VISUALS

Interactive World Simulator

A separately learned visual simulator can supply demonstrations and rank robot policies, but its usefulness depends on the fidelity and coverage of its task-specific interaction data.

Read illustrated report →
Training and inference diagrams share a causal video-action transformer; a green feedback path returns real robot observations to its KV cache.
READING 305 · 6 ORIGINAL VISUALS

DreamZero

Joint video-action denoising makes a pretrained video model usable as a generalizing robot policy, while closed-loop feedback and Flash training trade computation against execution quality.

Read illustrated report →
Two-system diagram: a QwenVL planner encodes instruction and image history, while latent goals and paired RGB features condition a trajectory diffusion transformer.
READING 306 · 6 ORIGINAL VISUALS

DualVLN

DualVLN uses supervised pixel goals and learned latent queries to connect slow semantic planning with fast RGB-conditioned trajectories, improving navigation while leaving substantial dynamic-obstacle failures.

Read illustrated report →
Left-to-right pipeline from bread-placement instruction and RGB-D image through generated video, object mask, video depth and 2D tracks to 3D object flow and robot motion.
READING 307 · 6 ORIGINAL VISUALS

Dream2Flow

Dream2Flow turns generated object motion into robot tracking goals, gaining flexibility across embodiments while inheriting errors from video generation, 3D reconstruction and domain-specific control.

Read illustrated report →
Two parallel DiT stacks: multi-view frames enter a 1.6B video model, whose cross-attention arrow points toward a 160M action expert receiving robot state and noisy actions. Both stacks have 28 blocks.
READING 308 · 6 ORIGINAL VISUALS

Act2Goal

Act2Goal uses multi-scale visual-transition features to guide a separate action expert, improving goal-conditioned manipulation while leaving the exact temporal schedule and some adaptation details unresolved.

Read illustrated report →
Three panels show parallel video and action DiTs, separate modality projections feeding shared attention, and an image/text-conditioned action refinement transformer.
READING 309 · 6 ORIGINAL VISUALS

CoVAR

CoVAR couples pretrained video diffusion to a dedicated action diffusion branch, improving reported manipulation success while relying on additional action refinement for low-resolution LIBERO90.

Read illustrated report →
Pink language block, blue video block with observed and partially predicted frames, and green action decoder with separate repeat loops and video/action noise symbols.
READING 310 · 6 ORIGINAL VISUALS

mimic-video

A separate inverse-dynamics decoder turns video-model hidden states into robot actions, gaining decoder-data efficiency while leaving video adaptation costs and task-dependent inference choices unresolved.

Read illustrated report →
Observed images pass through frozen encoders into latent history tokens. Future tokens, dexterous actions and camera motion enter the DexWM predictor; predicted features branch to an image decoder and a keypoint predictor.
READING 311 · 6 ORIGINAL VISUALS

DexWM

DexWM learns finger-conditioned latent dynamics from human videos and uses them for robot trajectory search, but reliable manipulation still depends on simulation adaptation, planning initialization and contact handling.

Read illustrated report →
Human preference annotation interface with synchronized video, semantic mask, depth and 3D-box panels, a Behavioral Safety task label, an empty rating field and a rationale area.
READING 312 · 6 ORIGINAL VISUALS

WorldLens

WorldLens exposes the gap between convincing driving imagery and functional reliability by evaluating reconstruction, control, perception and human judgment alongside generation quality.

Read illustrated report →
Two panels show language/video encoding and robot action vectors, followed by one diffusion transformer jointly denoising future video latents and actions. A separate optional decoder converts latents into frames.
READING 313 · 5 ORIGINAL VISUALS

VideoVLA

Jointly predicting video latents and robot actions transfers useful video priors into manipulation, with strong transfer results but substantial execution gaps and inference latency.

Read illustrated report →
An initial scene and instruction feed a VLM planner and affordance grounder; masks, paths, and phase indices condition a video generator whose candidates return to a VLM judge.
READING 314 · 6 ORIGINAL VISUALS

MIND-V

MIND-V turns language into geometric video conditions and searches over generated futures, improving reported subtask success while adding inference cost and dependence on learned judges.

Read illustrated report →
Camera and LiDAR perception feed current and future scene branches, a trajectory decoder, and a language-conditioned evaluator whose scores select the final trajectory.
READING 315 · 6 ORIGINAL VISUALS

MindDrive

MindDrive predicts action-conditioned BEV futures to refine trajectory anchors, then uses a VLM scorer to select a plan, with stronger benchmark scores but an incomplete training recipe.

Read illustrated report →
Two observed-image streams pass through a 3D VAE and video transformer, then spatial and motion filters and separate formers; their tokens condition an action transformer alongside language and current-image tokens.
READING 316 · 6 ORIGINAL VISUALS

Video2Act

Video2Act filters observed-video features into spatial and motion conditions for a separate action diffusion policy, trading expensive perceptual refreshes for faster action-chunk generation.

Read illustrated report →
Three latent-action pipelines compare Genie, UniVLA and LatBot. LatBot encodes a language-conditioned frame sequence into yellow scene and blue motion tokens, then decodes a future frame and an action sequence.
READING 317 · 6 ORIGINAL VISUALS

LatBot

LatBot transfers a video teacher's scene and motion representations into a current-view policy, trading substantial annotated pretraining for stronger downstream manipulation.

Read illustrated report →
Dreamer diagram showing 3D VAE encoding, Gaussian noise entering the latent stream, position embeddings, sparse self- and cross-attention, four FFN experts, T5 text conditioning and VAE decoding.
READING 318 · 6 ORIGINAL VISUALS

GigaWorld-0

GigaWorld-0 expands VLA supervision through controllable video generation and calibrated 3D simulation, but the supplied experiments establish video quality more directly than downstream robot improvement.

Read illustrated report →
Newt receives language, state and optional image embeddings, outputs planned actions, and trains a latent rollout against stop-gradient future encodings.
READING 319 · 6 ORIGINAL VISUALS

Newt

A demonstration-pretrained latent world model improves one controller across 200 tasks, while trailing specialists and showing uneven performance without feedback.

Read illustrated report →
Two training branches share RynnVLA-002: language, state and images feed discrete or continuous action outputs on the left; images and actions feed generated images on the right.
READING 320 · 6 ORIGINAL VISUALS

RynnVLA-002

A shared action-and-image predictor improves robot learning through joint supervision, while a separate continuous head makes action chunks practical without imagined-image rollouts during control.

Read illustrated report →
Two-stage AINA diagram: human video supplies tracked object points and stereo depth; point histories enter Vector Neuron MLPs, a transformer and an MLP, with fingertip outputs passed to inverse kinematics.
READING 321 · 6 ORIGINAL VISUALS

AINA

AINA transfers human demonstrations to a dexterous robot through aligned 3D point trajectories, reducing dependence on robot training data while retaining scene-calibration and contact-control constraints.

Read illustrated report →
Two separate vision-language models: a lower value head feeds an advantage calculation and A>epsilon binarization into an upper VLA, which outputs language, discrete actions, and flow-generated continuous actions.
READING 322 · 6 ORIGINAL VISUALS

π*0.6 / RECAP

RECAP turns experience into binary advantage labels that improve a flow-matching robot policy, while relying on human feedback and a critic whose calibration and data recipe require care.

Read illustrated report →
Observation tokens pass through self-attention, cross-attention and an MLP to a decoded next observation. Fourier-mapped actions join time embeddings and modulate the transformer. The legend distinguishes pretrained, newly trained and frozen modules.
READING 323 · 6 ORIGINAL VISUALS

Policy Evaluation with Video World Models

Action-conditioned video rollouts can rank robot policies, but their usefulness depends on behavioral coverage, faithful dynamics and reliable success judgments.

Read illustrated report →
Diagram linking an initial car image to a vision encoder, autoregressive world states, language actions and a video decoder, with history connections.
READING 324 · 6 ORIGINAL VISUALS

PAN

PAN combines language-conditioned latent prediction with causal video diffusion to support iterative simulation, while its evidence leaves physical validity and individual component benefits unresolved.

Read illustrated report →
DUST diagram: instruction and current image enter a VLM; robot state/actions and future-image features enter MMDiT, followed by separate action and vision DiTs and decoders.
READING 325 · 6 ORIGINAL VISUALS

DUST

DUST lets action and future-vision streams exchange information while retaining separate denoising dynamics, improving manipulation with an optional latency cost for extra vision refinement.

Read illustrated report →
A frozen vision-language model supplies a KV cache to a trainable action expert, whose velocity and noise determine Gaussian transitions from initial noise to an action.
READING 326 · 6 ORIGINAL VISUALS

πRL

Tractable stochastic denoising enables online RL for flow-based robot policies, improving familiar-task execution while leaving new-task generalization limited.

Read illustrated report →
Cosmos-Reason1 layer features feed a projection and text embedding, which enters DiT cross-attention. Noise or prefix-conditioned video tokens enter from above; timestep embeddings modulate scale, shift and gates. A dashed vision branch is also drawn.
READING 327 · 6 ORIGINAL VISUALS

Cosmos-Predict2.5 & Transfer2.5

A shared video generator becomes useful for Physical AI through specialized conditioning and synthetic observations, while most reported gains still concern video prediction rather than closed-loop control.

Read illustrated report →
Architecture with RGBD, trajectory, prompt, state and discrete-action tokens feeding a vision-language expert, alongside an action expert receiving noise and producing robot actions.
READING 328 · 5 ORIGINAL VISUALS

GigaBrain-0

GigaBrain-0 turns world-model generation into diverse VLA training data, improving tested visual robustness while leaving the full generation and deployment recipes underspecified.

Read illustrated report →
Two-ring benchmark composition chart with seven inner perturbation categories and finer outer components. The original inner Layout label incorrectly reads 1599.
READING 329 · 6 ORIGINAL VISUALS

LIBERO-Plus

LIBERO-Plus exposes how simulation policies fail under controlled shifts, while augmentation improves scores without establishing robust instruction use or initial-state control.

Read illustrated report →
Random hand-position and body-height commands generate an offline dataset. An encoder–decoder and recurrent latent chain feed predicted-latent, termination and value heads, with a legend identifying each component.
READING 330 · 6 ORIGINAL VISUALS

Ego-Vision Contact Planning

A short-horizon latent planner turns offline random humanoid interactions into purposeful contact, trading greedy value selection for a biased but potentially less variable sequence score.

Read illustrated report →
Two rows compare video-frame forecasting and token-grid forecasting. Both move from context plus states to generation, then compare predictions with targets.
READING 331 · 6 ORIGINAL VISUALS

Revontuli: Two Humanoid World Models

State-conditioned video generation and autoregressive token prediction win two forecasting tracks, while their inference choices expose a gap between benchmark scores and useful imagined futures.

Read illustrated report →
Three alignment branches convert egocentric video and hand trajectories into synthesized robot video/action pairs, which join real robot data for VLA training.
READING 332 · 6 ORIGINAL VISUALS

MimicDreamer

MimicDreamer turns human demonstrations into robot-domain training pairs through stabilization, IK and conditional video synthesis, improving a separate VLA while retaining dependence on robot data and calibration.

Read illustrated report →
An initial image and instruction feed separate pixel and latent world-model branches. Their features enter a modulation block that conditions noisy-action denoising and predicted actions.
READING 333 · 6 ORIGINAL VISUALS

MoWM

MoWM adds predicted latent dynamics to detailed video-model features before decoding actions, improving CALVIN chain completion while leaving the cause of that improvement only partly isolated.

Read illustrated report →
Four successive video chunks labeled 32, 8, 16 and 24 frames, with motion-threshold crossings or an effector-state change shown beneath them.
READING 334 · 6 ORIGINAL VISUALS

LongScape

LongScape turns action-derived chunk lengths into a choice among four video denoisers, improving reported long-horizon video quality while leaving routing causality and execution utility untested.

Read illustrated report →
Two-stage diagram: demonstrations initialize a policy; offline images and actions train diffusion dynamics and a reward classifier. In stage two the frozen world model produces transitions for PPO updates to separate policy and value networks.
READING 335 · 6 ORIGINAL VISUALS

World4RL

World4RL refines a separate robot policy inside a frozen diffusion simulator, gaining practice without further robot interaction while remaining vulnerable to errors in learned dynamics and rewards.

Read illustrated report →
Two panels show an image-and-language policy predicting yellow latent-action chunks for a world model during pretraining, then a policy taking image, language and robot state to output robot-action chunks during finetuning.
READING 336 · 6 ORIGINAL VISUALS

LAWM

LAWM uses a temporary world model to teach an imitation policy latent action chunks from video, trading extra pretraining and imperfect visual supervision for stronger downstream robot adaptation.

Read illustrated report →
Spatial plan table at left feeds a video diffusion model in the center; a separate flow-guided action diffusion model at right acts on the environment, with spatial feedback returning along the bottom.
READING 337 · 6 ORIGINAL VISUALS

Spatial Policy

Spatial Policy couples geometry-conditioned video planning to a separate action diffusion policy and feedback loop, improving simulated control while depending on accurate spatial inputs and several incompletely specified interfaces.

Read illustrated report →
Left-to-right architecture with layout, reference, and text inputs; shared temporal and multi-view layers; cross-modal interaction and projection heads; and RGB, depth, and semantic video outputs.
READING 338 · 6 ORIGINAL VISUALS

MoVieDrive

MoVieDrive couples shared geometric video modeling with modality-specific attention to generate RGB, depth, and semantics together, improving auxiliary outputs at a measurable RGB-fidelity cost.

Read illustrated report →
Three training paradigms: demonstration supervision, simulator action/reward feedback, and a separate reward world model receiving real sensor data and policy actions.
READING 339 · 5 ORIGINAL VISUALS

IRL-VLA

IRL-VLA replaces simulator reward queries during policy refinement with a learned trajectory scorer, yielding a small aggregate planning gain while shifting the balance between progress, safety and comfort.

Read illustrated report →
Multi-view images and an instruction enter a visual transformer; horizontal arrows send intermediate visual features into a separate action transformer initialized with noise, whose output leads to robot trajectories.
READING 340 · 6 ORIGINAL VISUALS

Genie Envisioner

Genie Envisioner turns robotic video priors into fast action decoding and action-conditioned simulation, but its transfer requires supervision and its simulation evidence remains based on video metrics.

Read illustrated report →
A video U-Net sends hidden features downward through a CNN video adapter into a separate action U-Net; video noise and action noise enter distinct branches.
READING 341 · 6 ORIGINAL VISUALS

Video Policy

A video diffusion model can supply transferable manipulation features to a separate action denoiser, but the evidence depends on task-video adaptation and comes with costly inference.

Read illustrated report →
Vidar diagram with embodied pre-training and target fine-tuning data, a unified video/language observation space, a mask-and-regression inset, and a left-to-right video-to-action pipeline.
READING 342 · 6 ORIGINAL VISUALS

Vidar

Vidar transfers a video prior through a sparse inverse-dynamics adapter, gaining low-demonstration manipulation performance while depending on visible robot state and costly video generation.

Read illustrated report →
StreamVLN diagram with an RGB timeline, vision encoder and projector, one language model, cached states, action tokens, and temporal then voxel-based pruning of inactive windows.
READING 343 · 6 ORIGINAL VISUALS

StreamVLN

StreamVLN makes navigation dialogue efficient by reusing recent KV states and compressing older visual history, trading complete context for bounded working context and occasional window-transition latency.

Read illustrated report →
An observation feeds an encoder and belief state. A nested loop links proposed actions with simulated future beliefs, while an actual action returns to the environment. A goal and critic sit to the right.
READING 344 · 6 ORIGINAL VISUALS

Critique of World Model

GLP proposes hierarchical discrete and continuous simulation grounded by observation reconstruction, trading a simpler latent objective for a richer but incompletely specified world-model and agent system.

Read illustrated report →
Language, state and image encoders feed a shared language-model backbone with dream and action queries. Colored world-prediction branches appear inside a dashed training-only box; an action diffusion transformer appears on the right.
READING 345 · 6 ORIGINAL VISUALS

DreamVLA

DreamVLA improves manipulation by supervising selective future knowledge inside a shared backbone, but its action-conditioning mechanism and reproduction settings retain unresolved reporting ambiguities.

Read illustrated report →
Two frozen visual tokenizers feed RGB and depth Transformer branches. Actions enter both branches, depth features feed upward into RGB, and sampled point tracks regularize predicted RGB tokens.
READING 346 · 6 ORIGINAL VISUALS

RoboScape

Depth feedback and tracked-token supervision improve action-conditioned robotic video prediction, but their benefits trade across metrics and do not yet establish physical robot transfer.

Read illustrated report →
Two task pathways share a WorldVLA backbone: text and images produce action tokens on the left; text, images and actions produce future-image tokens on the right.
READING 347 · 6 ORIGINAL VISUALS

WorldVLA

WorldVLA shares one autoregressive generator across policy and world prediction, while masking cross-action attention to make action chunks more reliable.

Read illustrated report →
Three panels show point tracking and FSQ reconstruction, image-and-language-conditioned autoregressive motion prediction, and a proprioception-aware cross-attention action head.
READING 348 · 6 ORIGINAL VISUALS

AMPLIFY

A discrete motion reference lets video learning and robot action learning scale separately, but accurate image-plane motion does not by itself establish controllable dynamics.

Read illustrated report →
Two training diagrams: masked-video feature prediction with an EMA target encoder on the left, and next-frame feature prediction conditioned on robot actions and poses with frozen encoders on the right.
READING 349 · 6 ORIGINAL VISUALS

V-JEPA 2

Video feature prediction can support robot planning after action-conditioned post-training, but the demonstrated transfer depends on image subgoals, short planning horizons and camera placement.

Read illustrated report →
Video frames enter a VAE and dense captions enter an LLM; interleaved shot tokens pass through alternating spatial and temporal DiT blocks. Enlarged blocks show text and vision branches only in the spatial block.
READING 350 · 6 ORIGINAL VISUALS

Seedance 1.0

Seedance combines shared text/image-conditioned video generation with a separate resolution refiner and extensive post-training, yielding strong overall preference results while leaving individual component gains difficult to isolate.

Read illustrated report →
Three panels show gripper exclusion and object tracking, a conditioned diffusion U-Net with frozen and trainable module markers, and ManiFlow-110k source proportions.
READING 351 · 6 ORIGINAL VISUALS

3DFlowAction

Predicting object motion supplies a transferable manipulation target, while rendered-goal checks and robot-specific optimization handle the gap between an imagined trajectory and executable actions.

Read illustrated report →
Realman dual-arm platform alongside two CogVideoX diffusion models; the first predicts future flow and sends it to the second, which predicts future RGB video.
READING 352 · 6 ORIGINAL VISUALS

CogRobot

CogRobot uses predicted optical flow to guide bimanual video plans, improving reported video quality and trained-task execution while retaining a separate action policy for each task.

Read illustrated report →
Metric depth and normal videos and a background reference pass through VAE encoders into concatenated video latents. Object images pass through CLIP and cross-attention. A VAE decoder produces a joint multi-view video.
READING 353 · 6 ORIGINAL VISUALS

RoboTransfer

RoboTransfer turns recorded geometry into appearance-diverse multi-view training videos, improving a separate imitation policy while leaving physical fidelity and evaluation uncertainty unresolved.

Read illustrated report →
Human, gripper and dexterous-hand videos supply prompts and target initial images to a video model; green and orange branches illustrate cross- and same-embodiment predictions, with internal features extracted below.
READING 354 · 6 ORIGINAL VISUALS

Human Video as a Robot Prompt

Cross-embodiment video features and retargeted human actions enable robot skill transfer, but the evaluated skills remain represented in human training data.

Read illustrated report →
Policy latents pass through a projection and join a text branch before DiT blocks; a separate image branch supplies visual context. Generated clips flow to a success detector.
READING 355 · 6 ORIGINAL VISUALS

WorldEval

WorldEval turns policy-internal action embeddings into videos whose judged outcomes can rank robot policies, but simulator familiarity and imperfect action fidelity constrain that proxy.

Read illustrated report →
An observation/action input block feeds a green MLP action module. Its latent-action branch conditions a gray video module; predicted videos and controls loop back as later inputs.
READING 356 · 6 ORIGINAL VISUALS

ProphetDWM

ProphetDWM links an MLP action forecaster to a diffusion video model through learned action features, improving reported nuScenes prediction metrics while leaving closed-loop driving and precise counterfactual dynamics untested.

Read illustrated report →
State, orange action tokens and green future tokens share DiT layers; current observations enter cross-attention, and future observation embeddings supply a separate alignment target.
READING 357 · 6 ORIGINAL VISUALS

FLARE

Training a shared action transformer to anticipate compact future observations improves manipulation, while leaving explicit planning and the reliability of its latent dynamics untested.

Read illustrated report →
EVAC diagram with reference-image and delta-action token branches feeding cross-attention, and action/ray maps plus memory and observations feeding the diffusion network.
READING 358 · 6 ORIGINAL VISUALS

EnerVerse-AC

EVAC predicts camera observations from supplied robot actions, enabling learned policy evaluation and synthetic data generation while leaving physical fidelity only partly tested.

Read illustrated report →
World initialization with a scene image, instruction and left/right trajectories feeds a video generator; generated frames feed scene, motion and semantic evaluation.
READING 359 · 5 ORIGINAL VISUALS

EWMBench

EWMBench separates scene stability, motion consistency, and semantic alignment to expose failures hidden by plausible robot videos, but its offline scores remain proxies whose calibration and reporting need careful scrutiny.

Read illustrated report →
Two training stages with blue DINOv2 frame features, yellow T5 language, green task-irrelevant codes and newly introduced red task-centric codes.
READING 360 · 6 ORIGINAL VISUALS

UniVLA

UniVLA learns task-oriented action tokens from videos to share policy knowledge across embodiments, while retaining supervised control adaptation and a fixed temporal bottleneck.

Read illustrated report →
Two-stage CLAM diagram: an inverse dynamics model outputs a continuous latent action to a forward model and an action decoder; a locked pretrained inverse model then supplies targets for a latent-action policy.
READING 361 · 6 ORIGINAL VISUALS

CLAM

CLAM makes observation-only robot imitation executable by jointly grounding continuous latent actions, but still depends on labeled transition coverage and task-specific robot demonstrations.

Read illustrated report →
Training pipeline with simulation and real data feeding replay, history and height encoders, action-conditioned forward recurrence, pose integration and a failure-probability head.
READING 362 · 6 ORIGINAL VISUALS

Perceptive FDM

Predicting platform-specific motion and failure risk makes MPPI navigation more effective on tested rough terrain, while retaining dependencies on training coverage and planner tuning.

Read illustrated report →
Four-panel pipeline: multiview rendering alignment, physics fitting through image loss, policy training in digital cousins, and real observation/action feedback.
READING 363 · 6 ORIGINAL VISUALS

PIN-WM

PIN-WM turns image-based rigid-body identification into a training environment for a separate robot policy, then uses nearby physics and appearance variations to improve transfer.

Read illustrated report →
Three panels show UWM training, action sampling with future observations masked by noise, and inverse-dynamics action sampling conditioned on clean future observations.
READING 364 · 6 ORIGINAL VISUALS

Unified World Models

Independent action and image noise levels let one transformer share dynamics and policy learning, improving selected robot-control results while leaving reliable visual planning unproven.

Read illustrated report →
CoT-VLA combines prior multimodal training, robot demonstrations and action-less video; blue goal-image outputs precede yellow action outputs, with a dotted execution-to-observation loop.
READING 365 · 6 ORIGINAL VISUALS

CoT-VLA

A shared multimodal model predicts an image of the desired future before generating robot actions, improving several manipulation results at the cost of slow image generation and delayed feedback.

Read illustrated report →
World-model diagram linking five encoded camera views, spatial and temporal attention, cross-attention and adaptive layer normalization to decoded video, with training-task and rollout insets.
READING 366 · 5 ORIGINAL VISUALS

GAIA-2

GAIA-2 combines compact video latents with structured conditioning to generate controllable surround-view driving scenes, while its evidence leaves physical validity and downstream driving benefit unresolved.

Read illustrated report →
A left-to-right video autoencoder diagram with one spatial and two spatial-temporal downsampling blocks, a compact latent space, and mirrored decoding blocks.
READING 367 · 6 ORIGINAL VISUALS

Wan

Wan makes video diffusion practical through causal latent compression, progressive image-video training and reusable conditioning, while its evaluation establishes media-generation quality rather than an executed-action model.

Read illustrated report →
A red curve for a 2B model with L:G=5:1 and sw=1024 lies increasingly below a yellow global-only curve as context grows from 1K to 128K; the vertical axis is KV-cache memory in MB.
READING 368 · 6 ORIGINAL VISUALS

Gemma 3

Gemma 3 expands visual and language capabilities by compressing image representations and restricting most attention to local windows, while long-context task reliability and full training reproducibility remain limited.

Read illustrated report →
Eagle-2 vision-language features enter DiT cross-attention while embodiment-specific state and noised-action encoders feed the action pathway; a decoder returns a chunk through a K-iteration sampling loop.
READING 369 · 5 ORIGINAL VISUALS

GR00T N1

GR00T N1 turns heterogeneous robot and video experience into an embodiment-aware action policy, improving tabletop adaptation while leaving broad humanoid autonomy untested.

Read illustrated report →
A diffusion-transformer backbone receives weighted residuals from multiple three-block modality branches; the output is a noise prediction.
READING 370 · 6 ORIGINAL VISUALS

Cosmos-Transfer1

Cosmos-Transfer1 combines independently trained video-control branches through spatial and temporal weights, trading appearance fidelity against variation in selected regions.

Read illustrated report →
Three panels show camera reconstruction and recurrent dynamics, goal-conditioned actor-critic training with plan recognition and proposal, and language-guided action inference with camera feedback.
READING 371 · 5 ORIGINAL VISUALS

LUMOS

LUMOS learns language-guided manipulation by practicing against expert latent trajectories inside frozen learned dynamics, gaining longer task chains while inheriting the world model’s errors.

Read illustrated report →
Historical image tokens, repeated action chunks, and masked future-image tokens enter a shared Transformer; its joint latents feed separate video and action diffusion heads.
READING 372 · 6 ORIGINAL VISUALS

UVA

UVA shares a video–action representation but separates diffusion decoding, making visual supervision compatible with fast policy inference while leaving task-specific performance and latency tradeoffs.

Read illustrated report →
Three camera views, robot joint state, and a language instruction enter a Llama-2 policy; empty action embeddings produce a parallel 25-step chunk of 14-dimensional joint targets.
READING 373 · 6 ORIGINAL VISUALS

OpenVLA-OFT

OpenVLA-OFT turns an autoregressive VLA into a fast continuous action-chunk policy, while OFT+ uses language-conditioned vision to improve instruction following in bimanual tasks.

Read illustrated report →
Left: noisy six-waypoint actions pass through an encoder, a command-conditioned transformer receiving VaViM features, and a decoder with an iterative feedback arrow. Right: query-by-key blocks separate causal visual attention from within-time action attention.
READING 374 · 6 ORIGINAL VISUALS

VaViM & VaVAM

Video pre-training supports trajectory imitation, but stronger synthesis and closer expert matching do not guarantee safer closed-loop driving.

Read illustrated report →
Survey map with ecosystem at the top, five prediction spaces beside four application families, and performance tasks and future directions below.
READING 375 · 6 ORIGINAL VISUALS

Driving World Models: A Survey

The survey separates what a driving world model predicts from how a driving system uses it, revealing why better scene generation alone cannot establish reliable planning.

Read illustrated report →
Three panels show VQGAN image compression, observation-conditioned prediction and decoding of future latents, and a repeated denoising U-Net with visual cross-attention inputs.
READING 376 · 6 ORIGINAL VISUALS

VILP

VILP makes video-guided imitation learning faster by generating compressed future frames and translating them into actions with a separate policy, while retaining task-specific data and control limitations.

Read illustrated report →
Two-stage architecture: an LDM compresses Go and robot clips into latent codes; an autoregressive transformer predicts interleaved codes and image tokens. A causal feature encoder, progressively masked queries, quantizer and decoder appear at right.
READING 377 · 6 ORIGINAL VISUALS

VideoWorld

Predicting compact multi-step dynamics codes makes video-based task learning more effective, while robot execution still requires a separately supervised action decoder.

Read illustrated report →
Five numbered panels show normalized action traces, DCT frequency components, a sparse coefficient matrix, frequency-first flattening and BPE compression into action tokens.
READING 378 · 6 ORIGINAL VISUALS

FAST

Compressing action chunks into quantized frequency coefficients and BPE tokens makes autoregressive robot policies easier to train, but the resulting decoder remains slower than the compared flow-matching policy at inference.

Read illustrated report →
Four connected RSPA modules: LLM-generated environment rewards, masked multi-view key-horizon reconstruction, recurrent latent dynamics and imagined policy learning.
READING 379 · 6 ORIGINAL VISUALS

RoboHorizon

RoboHorizon couples staged reward generation and key-horizon visual learning to an imagined-control policy, improving simulated task success while leaving reward correctness and several training details unresolved.

Read illustrated report →
Figure 2 shows randomly selected training frames, iterative prediction of latent chunks using retained observations and predictions, and ray-map-conditioned spatial and temporal diffusion blocks. Colored legends distinguish observed, predicted, noise and EOS latents.
READING 380 · 6 ORIGINAL VISUALS

EnerVerse

EnerVerse transfers a sparse-memory, multi-view video prior into fast action-chunk prediction, while leaving the causal contribution of its geometry and data-refinement components incompletely isolated.

Read illustrated report →
Original architecture diagram: current spectrogram passes through an audio encoder, latent flow matching and audio decoder to a future spectrogram; image and future audio feed the robot policy. An inset marks button pressing, rising pitch and button releasing, beside bottle-filling and piano examples.
READING 381 · 2 ORIGINAL VISUALS

Audio World Models

Predicting audio continuations gives a separate robot policy future context, but sparse comparisons leave the causal benefit and generalization uncertain.

Read illustrated report →
Pipeline from a user prompt and function library through an LLM, trajectories and an HDMap generator to camera-projected structure and CLIP-conditioned video generation.
READING 382 · 6 ORIGINAL VISUALS

DriveDreamer-2

Language specifies traffic, trajectory-conditioned diffusion supplies roads, and concatenated camera views support coherent video synthesis, with a tradeoff between visual conditioning and appearance diversity.

Read illustrated report →
Three rows show dough compression represented by particles, cloth folding represented by colored keypoints, and a falling block arrangement represented by discrete objects. Columns run from initial observation through representation and human prediction to subsequent observation.
READING 383 · 6 ORIGINAL VISUALS

Learned Dynamics for Manipulation

Choosing what to predict can make robot dynamics easier to learn, but stronger structure shifts difficulty into perception and constrains which interactions the model can represent.

Read illustrated report →
Four numbered panels show teleoperation-based video-model adaptation, image-and-language video generation, pseudo-action labeling, and policy training from the resulting dataset.
READING 384 · 6 ORIGINAL VISUALS

DreamGen

DreamGen converts adapted video priors into offline robot demonstrations, trading manual action collection for imperfect pseudo-labels and costly synthetic generation.

Read illustrated report →
Three stages show an encoder-decoder learning a discrete transition label from two images, a VLM predicting that label from an image and instruction, and LAPA predicting a robot action after fine-tuning.
READING 385 · 6 ORIGINAL VISUALS

LAPA

LAPA learns discrete visual-change targets for policy pretraining, improving transfer from actionless videos while leaving robot grounding and precise grasping dependent on supervised adaptation.

Read illustrated report →
Actor and learner processes communicate through policy parameters and two replay buffers; orange intervention transitions appear in both buffers, while blue policy transitions appear in the RL buffer.
READING 386 · 5 ORIGINAL VISUALS

HIL-SERL

Human corrections make real-world RL exploration useful, while reward-based updates turn those experiences into fast, reliable task-specific manipulation.

Read illustrated report →
Four flow diagrams contrast purple instantaneous tangent arrows with orange average-velocity arrows aligned with chords between times r and t.
READING 387 · 6 ORIGINAL VISUALS

MeanFlow

Learning an interval-average velocity turns trajectory integration into a single prediction, at the cost of a derivative-based training target and configuration-sensitive optimization.

Read illustrated report →
Original conceptual Figure 1 plots Number of tasks against Time with three uncalibrated curves and panels for traditional PhD-hours programming, fleet-model collected data, and fleet model plus Helix. The full source caption and adjacent motivation are retained.
READING 388 · 2 ORIGINAL VISUALS

Helix

Helix jointly learns a slow semantic conditioner and a fast upper-body control policy, with latency-aware training but no controlled evaluation isolating that design's benefits.

Read illustrated report →
Two game frames enter a latent action encoder. A sampled action and a direct connection from the earlier frame enter the decoder, which predicts the later frame.
READING 389 · 6 ORIGINAL VISUALS

AdaWorld

Pretraining a world model on continuous latent actions makes new control interfaces easier to learn, while leaving action selection to an external planner.

Read illustrated report →
Colored text, speech and image tokens pass through modality-specific projections and feed-forward blocks, sharing a central attention operation; adjacent panels contrast discrete prediction and continuous diffusion.
READING 390 · 6 ORIGINAL VISUALS

Mixture-of-Transformers

MoT separates modality-specific transformer weights while retaining shared attention, improving multimodal learning efficiency at the cost of extra stored parameters and implementation-sensitive overhead.

Read illustrated report →
System overview with IMU propagation, blue LiDAR processing, salmon visual processing and a shared adaptive voxel map.
READING 391 · 6 ORIGINAL VISUALS

FAST-LIVO2

FAST-LIVO2 uses LiDAR geometry to stabilize sparse visual alignment in a shared-map filter, trading dependence on calibrated sensors and local photometric assumptions for efficient odometry.

Read illustrated report →
Two parallel training diagrams. Text states and actions enter a language model; images and quantized actions enter a video model. Sample groups are decoded, compared with ground truth and scored through a dashed GRPO feedback loop.
READING 392 · 6 ORIGINAL VISUALS

RLVR-World

Rewarding decoded transition predictions improves separate language and video world models, but the gains remain specific to the chosen metrics, data and downstream evaluation.

Read illustrated report →
A stereo-image preprocessing branch feeds position, material and motion features into a state projector, Transformer encoder and state predictor; a future-stereo branch supplies ground-truth geometry for the hybrid loss.
READING 393 · 6 ORIGINAL VISUALS

ParticleFormer

Material-aware particle attention and hybrid geometric supervision improve the tested dynamics and control pipelines, while requiring scene-specific training and external perception.

Read illustrated report →
Two-stage architecture: discrete language and FAST outputs during pre-training; subtask prediction and a noise-conditioned continuous action expert during post-training and inference.
READING 394 · 6 ORIGINAL VISUALS

π0.5

A shared policy transfers heterogeneous supervision into real-home control through discrete pre-training and a continuous action expert, while the incremental value of explicit runtime hierarchy remains uncertain.

Read illustrated report →
Front-camera images pass through a VQ-VAE encoder and relative actions through binning. Blue image tokens and orange action triplets interleave around one DrivingGPT transformer, then separate into image decoding and action unbinning branches.
READING 395 · 6 ORIGINAL VISUALS

DrivingGPT

A shared causal transformer predicts interleaved visual and relative-motion tokens, improving the reported video and non-reactive planning benchmarks while leaving the causal value of joint training unresolved.

Read illustrated report →
Two training diagrams: observation encoders and decoders shape discrete recurrent states on the left; a single observed start state branches into imagined states with action, reward and value outputs on the right.
READING 396 · 6 ORIGINAL VISUALS

DreamerV3

DreamerV3 makes latent-imagination policy learning robust across diverse tasks through balanced losses and scale control, while retaining benchmark-specific resource and environment choices.

Read illustrated report →
Two rows compare simulated Tilt and Crawl scenes with red ground scandots against corresponding grayscale depth images. A green marker highlights the Tilt boundary.
READING 397 · 6 ORIGINAL VISUALS

WMP

WMP turns a sensory world model into recurrent context for a separate locomotion policy, improving difficult terrain traversal while trading perception-update speed against onboard computation.

Read illustrated report →
History images and a goal image pass through DINOv2. History features feed repeated action-conditioned dynamics blocks; a terminal predicted feature grid is compared with goal features. Orange arrows identify actions optimized at test time.
READING 398 · 6 ORIGINAL VISUALS

DINO-WM

Frozen spatial features let DINO-WM learn action-conditioned dynamics and plan toward goal images, but coverage of offline actions and the cost of test-time search remain limiting factors.

Read illustrated report →
Three scene rows compare real images with raw simulation, simulation plus real backgrounds, and simulation with both backgrounds and matched foreground textures.
READING 399 · 6 ORIGINAL VISUALS

SIMPLER

SIMPLER makes simulated evaluation useful for selecting real robot policies by matching visual inputs and calibrating action effects, while accepting that absolute success and difficult contact dynamics may still differ.

Read illustrated report →
Historical camera images and ego motions enter DCAE and a shared transformer; compact latent F feeds separate image and trajectory diffusion heads, with an action connection into image prediction and an autoregressive feedback loop.
READING 400 · 6 ORIGINAL VISUALS

Epona

Epona shares a causal history representation across separate trajectory and image diffusion heads, gaining flexible video rollouts and cheaper planner-only inference while still accumulating prediction error.

Read illustrated report →
Two time-shifted training windows show observation circles, GRU hidden-state squares, predicted observation feedback, action labels and observation losses.
READING 401 · 6 ORIGINAL VISUALS

RWM

RWM exposes a recurrent simulator to its own prediction errors during training, enabling long imagined PPO rollouts at the cost of sequential training and continued reliance on simulation.

Read illustrated report →
Bottom-to-top policy diagram showing shared image and proprioceptive branches, a robot wrist branch, context tokens, encoder trunk, action tokens and decoder, with separate OT and BC loss labels.
READING 402 · 6 ORIGINAL VISUALS

EgoBridge

Action-guided optimal transport improves human-to-robot policy transfer by aligning observation features during training, while relying on tracked human actions and task-specific correspondence.

Read illustrated report →
World4Drive pipeline: an intention encoder and a physical latent encoder feed trajectory generation; proposed trajectories condition future-latent prediction and a selector. A detached observed-future branch and expert trajectory provide training targets.
READING 403 · 6 ORIGINAL VISUALS

World4Drive

World4Drive ranks candidate driving trajectories through predicted latent futures, improving reported collision metrics while learning from future resemblance rather than an explicit safety objective.

Read illustrated report →
Three panels compare direct sensor-to-action driving, a VLM reasoning and multitask pathway with database fine-tuning, and an encoder-to-LLM/VLM-to-action-decoder pathway.
READING 404 · 5 ORIGINAL VISUALS

VLA4AD Survey

The survey organizes driving VLAs by where language enters action generation, while exposing the need to evaluate instruction following and executed driving under compatible protocols.

Read illustrated report →
Training diagram with Qwen2.5-VL-72B supplying reasoning annotations to Qwen2.5-VL-3B, and sampled fast/slow outputs scored in a GRPO loop.
READING 405 · 6 ORIGINAL VISUALS

AutoVLA

A shared language-and-motion decoder can improve driving plans and shorten reasoning, but its gains depend on data scale, decoding choices and benchmark conditions.

Read illustrated report →
Seer's multimodal encoder connects a foresight readout to an action readout, with separate future-image and action decoding branches.
READING 406 · 6 ORIGINAL VISUALS

Seer

Seer makes predicted visual latents available to an inverse-dynamics policy inside one transformer, improving task adaptation while leaving other-embodiment transfer much less reliable.

Read illustrated report →
Framework diagram showing robot and Internet data, three visual encoders and a language prompt entering a VLM, with robot state and noise entering a smaller action expert that outputs actions for multiple robot types.
READING 407 · 6 ORIGINAL VISUALS

π0

A VLM and a smaller flow-matching action expert turn diverse robot pre-training into reusable control, but dexterous task mastery still depends on post-training and execution design.

Read illustrated report →
Three-stage diagram taking a hammer photograph through geometry and texture generation, annotated tool axes, and task decomposition, constraints and key-pose calculation.
READING 408 · 6 ORIGINAL VISUALS

RoboTwin

RoboTwin uses annotated generative digital twins and language-generated robot programs to expand demonstration data, improving physical transfer while leaving coordinated dual-arm manipulation difficult.

Read illustrated report →
Five colored panels show before-and-after robot images for grasping, rotating, placing on a fixture, regrasping and inserting; downward arrows connect each operation’s initial and final states.
READING 409 · 6 ORIGINAL VISUALS

FMB

FMB makes functional assembly a modular physical benchmark, where useful grasps, task-specific sensing and human-assisted skill composition matter more than aggregate primitive success alone.

Read illustrated report →
Text and VAE video embeddings enter shared full-attention and feed-forward blocks, with separate text and vision scale, shift, gate and residual paths controlled by diffusion timestep t.
READING 410 · 6 ORIGINAL VISUALS

CogVideoX

CogVideoX combines temporal video compression, modality-specific normalization and full spatiotemporal attention to improve text-conditioned motion generation, at the cost of expensive attention and a substantial data pipeline.

Read illustrated report →
A diffusion-transformer diagram with frozen Qwen2.5 and SigLiP encoders, a temporal-token block, alternating conditioning arrows, a trainable coordinate decoder and a separate action expert.
READING 411 · 6 ORIGINAL VISUALS

A0

A0 makes manipulation transferable through sparse object-contact trajectories, while relying on a separate geometric execution stack to turn image coordinates into robot motion.

Read illustrated report →
Observation images and robot pose feed an action-sequence diffusion policy; a timeline separates prediction and subsequent replanning, beside CNN FiLM and transformer cross-attention denoisers.
READING 412 · 6 ORIGINAL VISUALS

Diffusion Policy

Denoising entire action sequences captures coherent alternatives, while executing only a short segment trades temporal consistency against feedback delay.

Read illustrated report →
One real camera view and three synthetic views sit above a trajectory diagram and the weighted two-stage scoring formula.
READING 413 · 6 ORIGINAL VISUALS

Pseudo-Simulation / NAVSIM v2

Pre-rendered future observations and endpoint-weighted scoring make driving evaluation sensitive to local deviations while retaining parallel planner queries, at the cost of incomplete interaction and per-scene reconstruction.

Read illustrated report →
Two-stage VPP diagram: CLIP and a video U-Net feed interpolated feature maps into spatial-temporal Video Former attention, then a DiT action denoiser.
READING 414 · 6 ORIGINAL VISUALS

Video Prediction Policy

VPP trades full video generation for one pass through a frozen video predictor, then learns a separate action policy from its predictive features.

Read illustrated report →
Student–teacher diagram showing observation and goal encoders, a history adaptation module, a dynamics decoder and FiLM, with action and next-state outputs and dashed teacher supervision links.
READING 415 · 6 ORIGINAL VISUALS

DyWA

DyWA uses history-conditioned joint action and task-state prediction to improve single-view non-prehensile manipulation, while relying on privileged simulation supervision and a small real-world evaluation.

Read illustrated report →
Three-part overview: observations and motion enter CDiT; two indoor predicted trajectories receive goal-image scores; outdoor sequences illustrate imagined motion in unknown environments.
READING 416 · 6 ORIGINAL VISUALS

NWM

A diffusion model can help select navigation actions by predicting their visual consequences, but accurate short-horizon ranking does not make unfamiliar-scene imagination a reliable map.

Read illustrated report →
Two vertical pipelines show human and robot hands converted into blue hand particles, red object particles and purple displacement arrows. Training produces object-state predictions; deployment compares primitive-action outcomes against a goal.
READING 417 · 5 ORIGINAL VISUALS

Cross-Embodiment Particle World Models

Particle displacement actions let one dynamics predictor serve different hands, while deployment still depends on known kinematics, reliable 3D perception and a restricted MPC search.

Read illustrated report →
Left-to-right GWM pipeline: optional second unposed image, Splatt3R reconstruction, 3D VAE compression, noisy latents, time and action conditioning, diffusion transformer and decoded future Gaussian splats. The action input is labeled Q and latent inputs KV.
READING 418 · 6 ORIGINAL VISUALS

GWM

GWM predicts action-conditioned futures in a compact Gaussian scene space, improving average policy performance while leaving exact conditioning and several evaluation comparisons unresolved.

Read illustrated report →
Five architecture panels show image and speed encoding, a Gaussian state-transition module, observation-conditioned estimation, observation-free prediction, and route-conditioned action decoding.
READING 419 · 6 ORIGINAL VISUALS

X-MOBILITY

A recurrent world model pretrained on random actions supplies a route-conditioned imitation policy with useful navigation state, while its forecasting and transfer evidence remain bounded by the tested scenes.

Read illustrated report →
One image branches into DINOv2 and SigLIP, a projector and language tokens feed Llama 2, and action tokens pass through a de-tokenizer to a robot arm.
READING 420 · 6 ORIGINAL VISUALS

OpenVLA

OpenVLA adapts a pretrained vision-language model into a direct action-token policy, gaining broad manipulation ability while remaining sensitive to training diversity, scoring rules and controller timing.

Read illustrated report →
Historical camera observations pass through a BEV encoder and conditioned memory to a world decoder; future occupancy and BEV features feed safety costs, candidate selection and a planner, whose trajectory returns through an action switch.
READING 421 · 6 ORIGINAL VISUALS

Drive-OccWorld

Drive-OccWorld couples an action-conditioned occupancy predictor with a cost-based planner, improving reported open-loop planning while leaving closed-loop safety and implementation reproducibility unresolved.

Read illustrated report →
Pipeline with RGB-D observation, language and proprioception; sparse-convolutional voxel features feed a multimodal action policy and an auxiliary task-labeled Gaussian model with sequential leader and follower deformation.
READING 422 · 5 ORIGINAL VISUALS

ManiGaussian++

A task-labeled Gaussian world model predicts stabilizing-then-acting effects to improve a separate bimanual policy, trading stronger reported multi-task success for calibrated multi-view supervision and unresolved implementation details.

Read illustrated report →
A flowchart divides edge-side collection and validity checks from cloud-side processing, quality checks, recovery annotation, delivery and model-training feedback.
READING 423 · 6 ORIGINAL VISUALS

AgiBot World & GO-1

AgiBot World couples standardized robot demonstrations with latent-action pretraining, improving selected manipulation scores while leaving the effects of data scale, quality and embodiment alignment partly entangled.

Read illustrated report →
Two panels show GPC-RANK evaluating several imagined action rollouts and GPC-OPT refining a single proposal through reward gradients; red stars mark selected outcomes.
READING 424 · 6 ORIGINAL VISUALS

GPC

GPC improves a frozen diffusion policy by ranking or refining its actions through a separate predictive world model, trading better manipulation outcomes for substantial inference cost.

Read illustrated report →
LAW architecture: current images produce visual latents and waypoints, both feed an action-aware latent predictor, and future encoded latents supervise predictions. Lower panels contrast perspective-view and BEV planners.
READING 425 · 6 ORIGINAL VISUALS

LAW

LAW improves a driving planner by predicting action-conditioned future features during training, trading explicit future-image generation for an auxiliary latent objective whose gains depend on target horizon and evaluation setting.

Read illustrated report →
Driving images, perception text and trajectories are reorganized into a time sequence, passed through modality tokenizers, and modeled by one next-token prediction block with a shifted target row.
READING 426 · 6 ORIGINAL VISUALS

Doe-1

Doe-1 learns images, descriptions and ego motion in one autoregressive stream, gaining several driving capabilities while leaving the safety and efficiency of driving with live feedback unestablished.

Read illustrated report →
GR-2 at the center of blue video-language pre-training examples and orange robot fine-tuning examples, connected by double-headed arrows.
READING 427 · 6 ORIGINAL VISUALS

GR-2

GR-2 transfers large-scale video prediction into a policy that generates future images and action trajectories, achieving strong manipulation results while leaving the causal role of visual planning incompletely tested.

Read illustrated report →
VQ-VAE encoders turn observations into token sequences interleaved with actions below a shared transformer. Separate observation, reward, termination and Q-value heads appear above it; a legend identifies the colored node types.
READING 428 · 6 ORIGINAL VISUALS

JOWA

Sharing a transformer between world prediction and conservative Q-learning improves aggregate offline Atari performance, while short imagined searches add return at an inference-time cost.

Read illustrated report →
A scene image and task prompt feed a video model. Generated human frames and robot observation frames enter separate visual encoders, then a cross-attention action head. Dotted branches connect both streams to auxiliary point-track predictors.
READING 429 · 6 ORIGINAL VISUALS

Gen2Act

Gen2Act uses a generated human demonstration to guide a separate robot policy, gaining unseen-task coverage while remaining limited by both video plausibility and execution quality.

Read illustrated report →
Two-part architecture diagram: individual images enter a blue encoder and pink inverse dynamics model; transition latents join observation embeddings before a green forward model predicts future embeddings.
READING 430 · 6 ORIGINAL VISUALS

DynaMo

DynaMo uses compact latent transitions and future-embedding prediction to pretrain useful robot vision features, while action selection remains the job of a separately trained policy.

Read illustrated report →
Scene tokenizer above text, scene and action vocabularies feeding one OccLLaMA block; an attention matrix shows a blue scene block within gray causal attention.
READING 431 · 6 ORIGINAL VISUALS

OccLLaMA

A shared occupancy-language-action generator improves several driving benchmarks, but its spatial attention helps scene forecasting more consistently than trajectory planning.

Read illustrated report →
An image passes through SigLIP and a linear projection; tokenized question text joins the image-token sequence before the Gemma decoder generates an answer.
READING 432 · 5 ORIGINAL VISUALS

PaliGemma

PaliGemma turns pretrained vision and language components into a compact, adaptable VLM through joint input attention and broad training, with task-specific fine-tuning and resolution costs determining the final behavior.

Read illustrated report →
Video diffusion architecture with VAE-encoded gesture images and first frame, noisy frames, a zero-convolution conditioning branch, UNet, CLIP text/image embeddings, and a stage-specific color legend.
READING 433 · 6 ORIGINAL VISUALS

This&That

Language and pointing disambiguate a generated video plan, while a separate controller learns to follow that plan; the demonstrated execution benefit is confined to simulation.

Read illustrated report →
A left training branch connects stereo human demonstrations to video diffusion fine-tuning; a right inference branch connects a robot-scene image to generated video, CAD-based tracking, and robot actions.
READING 434 · 6 ORIGINAL VISUALS

Dreamitate

Dreamitate converts generated human tool-use videos into executable robot poses, gaining task-level generalization while relying on trackable rigid tools and slow open-loop inference.

Read illustrated report →
Instruction parsing produces pick orange, from bottom drawer and place on counter. Four conditioned noise-prediction branches, including a goal image, merge into a composed output video.
READING 435 · 5 ORIGINAL VISUALS

RoboDreamer

Composing phrase-conditioned video predictions improves alignment with unfamiliar robot instructions, while execution still depends on a separate controller and incomplete spatial information.

Read illustrated report →
PAD architecture with RGB, robot-pose and optional extra-modality columns passing upward through a shared diffusion transformer; an expanded block shows masked attention and frozen CLIP conditioning.
READING 436 · 6 ORIGINAL VISUALS

PAD

PAD couples future-image prediction and robot-pose generation in one denoiser, improving reported manipulation success while increasing inference cost.

Read illustrated report →
A road image with orange, green and red candidate paths, followed by BEV panels showing the human trajectory and candidate road/collision outcomes with ADE values.
READING 437 · 6 ORIGINAL VISUALS

NAVSIM

NAVSIM trades reactive feedback for inexpensive evaluation on real sensor observations, using short simulated trajectories and curated scenes to make planning scores more informative.

Read illustrated report →
Two-stage GR-1 overview with a modality legend, human-video pretraining on the left, copied weights and robot fine-tuning on the right, plus CALVIN and physical-robot examples.
READING 438 · 6 ORIGINAL VISUALS

GR-1

GR-1 uses a shared transformer to transfer human-video forecasting into robot action learning, with strong task-chain gains but incomplete evidence about which aspects of prediction cause those gains.

Read illustrated report →
Curated and uncurated image collections feed embedding, deduplication and retrieval stages; selected web images join the curated collection.
READING 439 · 6 ORIGINAL VISUALS

DINOv2

DINOv2 combines curated image retrieval with global and patch-level self-supervision to make frozen visual features broadly useful, while leaving task decoding and domain bias unresolved.

Read illustrated report →
Handheld and robot-mounted grippers flank fisheye camera views, with six labels identifying camera, lens, side mirrors, IMU tracking, width tracking and feasibility filtering.
READING 440 · 6 ORIGINAL VISUALS

UMI

UMI makes portable handheld demonstrations transferable by matching sensing, action coordinates and execution timing, while remaining limited by robot feasibility and the coverage of the collected data.

Read illustrated report →
Colored trajectory states and red action arrows evolve around a black disk; solid arrows connect action updates and hollow arrows connect state updates. The lower-left panel depicts the policy field.
READING 441 · 6 ORIGINAL VISUALS

PolyGRAD

PolyGRAD refines whole imagined trajectories using a denoiser and policy-score action updates, gaining a non-autoregressive sampler while making action calibration a central training constraint.

Read illustrated report →
Three-part benchmark overview: a Think2Drive expert and training set on the left, interactive driving scenes in the center, and schematic ability-assessment plots on the right.
READING 442 · 6 ORIGINAL VISUALS

Bench2Drive

Short, scenario-focused CARLA routes make driving failures easier to diagnose, while shared expert data improve comparability within the limits of simulated behavior and uneven coverage.

Read illustrated report →
Dataset comparison with duration, front-view frame counts, geographic coverage and sensor setup; the final rows show OpenDV-YouTube and the combined OpenDV-2K collection.
READING 443 · 6 ORIGINAL VISUALS

GenAD / OpenDV-2K

GenAD learns driving-video dynamics around a frozen image backbone, enabling controllable visual futures and a small planning head while retaining costly pretraining and unproven closed-loop performance.

Read illustrated report →
Three rows of linked carts contrast initial anchoring, measurement anchoring without an initial prior, and an entirely unanchored chain.
READING 444 · 5 ORIGINAL VISUALS

State Estimation — Chapters 3 and 8

Gaussian structure makes estimation exact and sparse, while rigid-body pose estimation preserves geometry through local updates whose approximations must remain explicit.

Read illustrated report →
Annotated photograph of a Franka Panda on a portable desk, identifying its gripper, two external stereo cameras, wrist camera, laptop and teleoperation headset.
READING 445 · 5 ORIGINAL VISUALS

DROID

DROID makes varied real-world demonstration collection practical with a shared portable robot platform, improving average co-trained policy success while leaving zero-shot deployment and the isolated effect of scene diversity unresolved.

Read illustrated report →
Three dog observations feed encoders and normalized latent grids. Action-labeled arrows connect successive grids, and each grid emits predicted action, reward, and value.
READING 446 · 6 ORIGINAL VISUALS

TD-MPC2

TD-MPC2 makes latent-space planning more robust across continuous-control tasks, while retaining an explicit planner and substantial dependence on task rewards and dataset coverage.

Read illustrated report →
Actor and learner nodes share a parallel-training block above a robot environment; actions flow down to the robot and transition tuples flow back up to the learner.
READING 447 · 6 ORIGINAL VISUALS

SERL

SERL turns demonstration-assisted RL into an effective real-robot training stack, while its integrated results leave the contribution of individual engineering choices unresolved.

Read illustrated report →
Two-panel diagram showing robot frames passing through a video tokenizer, padded history and noisy trajectory conditioning Open-Sora, and learned action queries feeding an adapter and action head.
READING 448 · 6 ORIGINAL VISUALS

VidMan

VidMan transfers video-diffusion features into a shared-backbone action policy, improving reported manipulation results while avoiding iterative video generation during control.

Read illustrated report →
Training diagram with a red generator, locked gray real-score network, blue fake-score network, and green discriminator sharing the fake network’s encoder.
READING 449 · 6 ORIGINAL VISUALS

DMD2

DMD2 accelerates image diffusion through better critic tracking, real-data supervision and student-generated training inputs, while retaining guidance, diversity and implementation tradeoffs.

Read illustrated report →
Eleven-line GKD pseudocode with a uniform draw, an on-policy branch when u is at most lambda, a fixed-data alternative and a student-gradient update.
READING 450 · 6 ORIGINAL VISUALS

GKD

GKD teaches a smaller language model on the prefixes it generates, improving distillation while adding sampling cost and a task-dependent choice of divergence.

Read illustrated report →
Three columns compare guiding, contact and minimum-time rewards, with fixed versus success-terminated episodes.
READING 451 · 6 ORIGINAL VISUALS

Revisiting Sparse Rewards

Success-terminated constant rewards can improve final reaching behavior, while reliable learning still depends on how often exploration encounters the goal.

Read illustrated report →
Repeated time columns contain scene tokens and a purple ego token, spatial aggregation blocks, shared temporal causal attention, separate occupancy and displacement decoders, and red dashed feedback arrows.
READING 452 · 6 ORIGINAL VISUALS

OccWorld

OccWorld jointly predicts semantic occupancy and ego motion through shared world tokens, but its strongest results depend on oracle occupancy and deteriorate with forecast horizon.

Read illustrated report →
Ten kitchen floor plans in two rows, from one-wall and galley layouts to U-shaped, G-shaped and wraparound rooms.
READING 453 · 6 ORIGINAL VISUALS

RoboCasa

RoboCasa turns diverse simulated kitchens and human demonstrations into larger imitation-learning datasets, improving average atomic-task success while leaving composite behavior and real-world reliability unresolved.

Read illustrated report →
DriveDreamer diagram showing text, reference image, initial HDMap, boxes, and actions feeding ActionFormer and Auto-DM, with separate video and action outputs.
READING 454 · 6 ORIGINAL VISUALS

DriveDreamer

DriveDreamer forecasts structural scene features from actions, renders them with a shared diffusion backbone, and decodes future actions, trading strong conditional generation for dependence on structured inputs and incompletely specified evaluation details.

Read illustrated report →
Noise enters the one-step generator. Its output branches to paired regression and to noisy-image processing by a locked real-score network and a trainable fake-score network. A red score-difference path feeds the generator.
READING 455 · 6 ORIGINAL VISUALS

DMD

DMD trades expensive distillation training for one-pass image generation by combining a real-versus-generated score gradient with teacher-pair regression.

Read illustrated report →
Training frames split into a blue video-tokenizer branch and a yellow latent-action branch; both feed an orange dynamics model that predicts future frames.
READING 456 · 6 ORIGINAL VISUALS

Genie

Genie turns unlabeled video into an interactive world model by learning a small latent action vocabulary, trading broad video supervision for uncertain action semantics and short-horizon consistency.

Read illustrated report →
Three columns depict reverse diffusion from noisy to clean observations, followed by policy actions; past images and actions condition each column.
READING 457 · 6 ORIGINAL VISUALS

DIAMOND

EDM diffusion preserves useful Atari visual details with a short sampling schedule, enabling effective imagined policy training while retaining separate reward and action networks.

Read illustrated report →
Three labeled pixel observations compare Crafter, Craftax-Classic and Craftax, each showing a local terrain view above player statistics and inventory.
READING 458 · 6 ORIGINAL VISUALS

Craftax

Craftax makes long-budget exploration experiments inexpensive, but its tested policies still struggle to turn local reward collection into deeper progression.

Read illustrated report →
Three labeled photographs show an ARL Warthog, an ARL Jackal carrying a tag board, and an NYU ARPL aerial robot with a forward-facing camera.
READING 459 · 6 ORIGINAL VISUALS

CoPeD

CoPeD makes complementary real-world air-ground observations available for collaborative perception, trading inexpensive automatic annotation for unresolved label accuracy and largely qualitative validation.

Read illustrated report →
Architecture diagram with green task tokens, blue observation tokens, purple readouts and action heads, followed by a finetuning layout adding observation and action interfaces.
READING 460 · 6 ORIGINAL VISUALS

Octo

Octo makes a shared robot policy adaptable through modular token interfaces and diffusion action decoding, but its strongest transfer results still require target demonstrations and full-model finetuning.

Read illustrated report →
Two camera views enter a ViT and resampler, language enters a fusion decoder, pooled multimodal features enter an LSTM, and the head outputs actions. The legend identifies token colors and frozen modules.
READING 461 · 6 ORIGINAL VISUALS

RoboFlamingo

RoboFlamingo turns single-step vision-language features into a recurrent manipulation policy, gaining CALVIN performance while risking loss of the backbone's general vision-language abilities.

Read illustrated report →
Four colored design-axis boxes surround an image encoder, projection block and language model. Arrows run from visual features to projection and upward to the LM; snowflake and flame symbols identify frozen and trainable components.
READING 462 · 6 ORIGINAL VISUALS

Prismatic VLMs

Prismatic improves image-conditioned language models through simpler optimization and richer patch features, while exposing task tradeoffs that aggregate scores can hide.

Read illustrated report →
Five panels show dataset counts, scene and trajectory shares by robot, frequent skill labels, and frequent objects grouped into categories.
READING 463 · 6 ORIGINAL VISUALS

Open X-Embodiment / RT-X

Pooling robot demonstrations can transfer skills across embodiments through direct action policies, but the gain depends on model configuration and task distribution.

Read illustrated report →
A generated-video input passes through I3D and two Transformer encoder layers, branching into a pooled action classifier and an attention-equipped GRU trajectory decoder.
READING 464 · 6 ORIGINAL VISUALS

ACT-Bench

ACT-Bench measures action fidelity through a learned video-to-motion evaluator, revealing TERRA's higher aggregate controllability while leaving evaluator transfer and comparison fairness unresolved.

Read illustrated report →
Four source-video strips with side-by-side checkmarks for cut detection without and with a cascade; gradual transitions are missed without the cascade.
READING 465 · 6 ORIGINAL VISUALS

Stable Video Diffusion

Curated video pretraining gives a latent diffusion backbone a transferable motion representation, but its strongest evidence remains short-video preference and image-based multi-view evaluation.

Read illustrated report →
Interleaved image and action tokens feed an MLLM; its current action feeds diffusion, whose generated next image returns along green arrows. Dashed arrows mark historical training outputs.
READING 466 · 6 ORIGINAL VISUALS

ADriver-I

ADriver-I links language-based control prediction to action-conditioned video generation, enabling recurrent imagined driving while leaving physical fidelity and long-term reliability unverified.

Read illustrated report →
Image, action and text encoders feed interleaved token groups into an autoregressive world model; predicted image tokens then pass to a video decoder.
READING 467 · 6 ORIGINAL VISUALS

GAIA-1

GAIA-1 separates action-conditioned image-token prediction from diffusion video rendering, enabling controllable driving scenarios while leaving physical accuracy and downstream driving benefits unmeasured.

Read illustrated report →
Three training stages show image features flowing upward through ViT and cross-attention to QwenLM, with learnable queries entering each adapter and snowflake or flame symbols marking frozen or trainable components.
READING 468 · 6 ORIGINAL VISUALS

Qwen-VL

Qwen-VL combines spatially informed visual compression with staged language supervision to read, describe and localize image content, while its ablations leave downstream causal benefits only partly established.

Read illustrated report →
Seven pretraining sources with sampling proportions, epochs and disk sizes; CommonCrawl has the largest sampling share.
READING 469 · 6 ORIGINAL VISUALS

LLaMA

LLaMA spends more training tokens on compact causal language models to improve capability at a chosen inference size, while leaving the contribution of individual architectural choices unresolved.

Read illustrated report →
UniPi diagram: initial observation and text, tiled image context, video diffusion, sparse frames, temporal super-resolution, dense frames, then inverse dynamics and robot actions.
READING 470 · 6 ORIGINAL VISUALS

UniPi

UniPi shares planning through text-conditioned videos, but translating those plans into reliable behavior still depends on a separate action model and executable image trajectories.

Read illustrated report →
Pipeline from sparse SfM points through initialization to 3D Gaussians, camera projection, differentiable tile rasterization and an image, with returning gradient arrows and adaptive density control.
READING 471 · 6 ORIGINAL VISUALS

3D Gaussian Splatting

A scene-specific set of anisotropic Gaussians makes high-quality radiance fields fast to render, at the cost of substantial memory and approximate visibility.

Read illustrated report →
PaLM receives interleaved orange language, green observation and blue image embeddings, generates text and passes it leftward to a control block; surrounding panels illustrate robot and language tasks.
READING 472 · 6 ORIGINAL VISUALS

PaLM-E

PaLM-E learns to turn continuous observations into language-conditioned plans, gaining from broad multimodal training while relying on separate robot controllers and facing a tradeoff between adaptation and language retention.

Read illustrated report →
RT-1 architecture: a language embedding conditions six image encoders through FiLM, TokenLearner reduces their spatial tokens, and a Transformer outputs mode, arm and base commands.
READING 473 · 6 ORIGINAL VISUALS

RT-1

RT-1 combines early language-conditioned visual compression with discrete action prediction to learn broad kitchen manipulation skills, trading limited motion novelty for practical feedback control.

Read illustrated report →
Three panels show green self-penetration sampling spheres, contact candidates on an open ShadowHand, and initial hands surrounding an object's translucent blue inflated convex hull.
READING 474 · 6 ORIGINAL VISUALS

DexGraspNet

DexGraspNet scales geometry-based grasp optimization and simulation filtering into useful training data, while leaving precision grasping and the balance between pose diversity and stability unresolved.

Read illustrated report →
A WidowX 250 arm between two movable cameras, with a fixed depth-camera view marked behind the workspace.
READING 475 · 6 ORIGINAL VISUALS

BridgeData V2

A shared manipulation dataset supports several action-learning methods and cross-lab reuse, while its single robot type, uneven coverage and limited evaluations bound the generalization claim.

Read illustrated report →
Two linked pipelines turn RGBD subviews into waypoint candidates and use point-cloud projection plus a generator to attach imagined panoramic observations to a hypothetical environment graph.
READING 476 · 6 ORIGINAL VISUALS

DREAMWALKER

A learned panoramic world model makes short-horizon waypoint search useful for continuous navigation, while prediction errors and goal-distance estimation limit the benefit of imagining further.

Read illustrated report →
Bottom-to-top latent denoising pipeline and three transformer block diagrams, with adaLN-Zero highlighted and conditioning-dependent scales drawn before residual additions.
READING 477 · 5 ORIGINAL VISUALS

DiT

DiT makes a transformer the latent-diffusion denoiser, improving ImageNet generation through adaptive conditioning and increased model computation while retaining a pretrained convolutional VAE.

Read illustrated report →
Four panels show crossing linear interpolations, rewired flow paths, interpolated new endpoints and a straighter second flow; purple and red dots mark the two distributions.
READING 478 · 6 ORIGINAL VISUALS

Rectified Flow

Learning from straight interpolations and then reflowing generated endpoint pairs makes image transport usable with very few neural evaluations, at the cost of additional training and possible full-solver quality loss.

Read illustrated report →
Source demonstrations are arranged by subtask; selected segments feed a loop that transforms a trajectory, interpolates to its start, executes it and observes the new scene.
READING 479 · 6 ORIGINAL VISUALS

MimicGen

Object-relative replay can turn a few human demonstrations into useful policy-training data, but collection yield and downstream learning quality respond differently to noise, interpolation and scene coverage.

Read illustrated report →
Multi-scale image features and positional embeddings feed an encoder. Query selection initializes decoder anchors, encoder outputs provide keys and values, and separate learned content queries and noisy ground-truth queries enter the decoder.
READING 480 · 6 ORIGINAL VISUALS

DINO

DINO improves object detection by choosing image-specific anchor positions, teaching noisy anchors when to reject objects, and sharing box supervision across adjacent decoder layers.

Read illustrated report →
Web question-answer and robot-action examples feed a ViT and language-model policy; generated tokens are de-tokenized into robot commands, with example deployments at right.
READING 481 · 6 ORIGINAL VISUALS

RT-2

RT-2 turns language generation into direct robot control, gaining semantic generalization while remaining limited by demonstrated motions and costly inference.

Read illustrated report →
A left-to-right pipeline converts HM3D and Gibson scans into navigation graphs, repaired panoramas, sampled trajectories and speaker-generated instructions, then uses the resulting dataset for downstream navigation training.
READING 482 · 6 ORIGINAL VISUALS

ScaleVLN

ScaleVLN improves navigation by generating diverse training routes with traversable graphs, repaired observations and synthetic instructions, but the benefit depends on data quality and how training uses it.

Read illustrated report →
Two training diagrams: observation encoders and decoders accompany recurrent latent dynamics on the left; actor, reward and value outputs accompany imagined latent transitions on the right.
READING 483 · 6 ORIGINAL VISUALS

DayDreamer

A learned latent simulator trains an actor while robots collect real experience, enabling learning within hours under carefully engineered control interfaces.

Read illustrated report →
An initial image enters a green encoder; image tokens feed a blue Transformer that predicts rewards, termination and future tokens. Green decoders feed a purple policy, which supplies actions back to the token sequence.
READING 484 · 6 ORIGINAL VISUALS

IRIS

IRIS turns compressed Atari images into an autoregressive training environment for a separate policy, trading real interaction for simulated experience whose usefulness depends on visual fidelity and event coverage.

Read illustrated report →
Four image–text similarity grids show three devices computing local blocks, rotating text representations, accumulating losses and taking a final cross-device sum.
READING 485 · 6 ORIGINAL VISUALS

SigLIP

Independent image–text matching losses make distributed pretraining more memory-efficient, while the experiments show that larger batches alone are a poor substitute for an appropriate training recipe.

Read illustrated report →
Four task-suite branches surround a procedural generator. The original left branches associate Object with changed layouts and Spatial with changed objects; five research topics run along the bottom.
READING 486 · 6 ORIGINAL VISUALS

LIBERO

LIBERO makes lifelong robot learning measurable through controlled task changes, exposing the tension between acquiring new skills and retaining old ones.

Read illustrated report →
Top view of a car marking a front camera, central VLS128, two VLP16 LiDARs, OxTS navigation sensor, rear-axle and vehicle centers, and colored local coordinate axes.
READING 487 · 6 ORIGINAL VISUALS

ZOD

ZOD uses diverse, calibrated driving keyframes and separate temporal recordings to broaden perception research, while its baselines expose unresolved long-range and rare-class failures.

Read illustrated report →
Green probability-flow trajectories connect a multimodal data distribution on the left to noise on the right; red consistency-map arrows return three points to one shared origin.
READING 488 · 5 ORIGINAL VISUALS

Consistency Models

Learning a common endpoint for every point on a diffusion trajectory enables one-step image generation, while optional repeated denoising improves quality and enables editing.

Read illustrated report →
Three related tabletop robot scenes labeled Training, Env A, Env B and Env C point to a held-out Test scene labeled Env D; each scene includes an example language instruction.
READING 489 · 5 ORIGINAL VISUALS

CALVIN

CALVIN turns language-driven robot manipulation into a test of skill chaining and scene transfer, revealing that competent isolated actions can coexist with almost no five-instruction completion.

Read illustrated report →
TD-MPC feedback-loop schematic above six Humanoid and Dog learning curves comparing red TD-MPC, blue SAC and dashed simulator MPC.
READING 490 · 6 ORIGINAL VISUALS

TD-MPC

A reward-oriented latent model and terminal Q function make short-horizon continuous-action planning effective, with task-dependent costs in computation and transfer.

Read illustrated report →
A goal photograph, aerial exploration and discovered routes with map nodes, robot views, and separate training and unseen-test environment examples.
READING 491 · 6 ORIGINAL VISUALS

RECON

RECON combines compressed, context-conditioned goal sampling with topological memory to accelerate outdoor exploration and return navigation, while relying on heuristic reachability and online adaptation.

Read illustrated report →
Input x splits into a blue pretrained weight block and an orange A-to-B rank-r branch; their outputs are added to form h. A has Gaussian initialization and B starts at zero.
READING 492 · 6 ORIGINAL VISUALS

LoRA (arXiv v2)

LoRA makes task adaptation small by learning low-rank weight updates that can be merged for inference, while useful rank and update placement remain task-dependent.

Read illustrated report →
Five labeled benchmark panels grouped under Past, Present and Future: Episodic Memory, Hands & Objects, Audio-visual Diarization, Social Interaction and Forecasting.
READING 493 · 6 ORIGINAL VISUALS

Ego4D

Ego4D supplies diverse first-person experience and task-specific supervision, while uneven sensor coverage and separately documented baselines limit what this paper alone can establish.

Read illustrated report →
Two trajectories supply a state-action pair, a future positive and a random negative. State-action encoder phi and goal encoder psi feed a comparison at the top.
READING 494 · 6 ORIGINAL VISUALS

Contrastive RL

A contrastive future-state discriminator can train a goal-reaching actor, but its value interpretation depends on how trajectories and goals are sampled.

Read illustrated report →
Annotated photograph of Jackal on the left and Spot on the right, showing RGB cameras, Velodyne lidar, Jackal stereo and wheel odometry, and Spot body cameras, visual odometry and joint angles.
READING 495 · 6 ORIGINAL VISUALS

SCAND

SCAND turns natural robot teleoperation into supervision for social navigation, with promising Spot baselines whose evidence remains limited to path imitation and simple participant-rated encounters.

Read illustrated report →
An image encoder and decoder connect pixel space to latent space. A forward diffusion arrow leads to noise; a UNet reverses the process. A condition encoder feeds cross-attention or concatenation.
READING 496 · 5 ORIGINAL VISUALS

Latent Diffusion Models

Moving iterative diffusion into a mildly compressed image space reduces computation, while reconstruction quality and sampling guidance determine what fidelity and coverage can be retained.

Read illustrated report →
A satellite view of the Pittsburgh collection site with colored driving traces, surrounded by four sets of RGB images, height and color maps, IMU traces and shock traces.
READING 497 · 6 ORIGINAL VISUALS

TartanDrive

TartanDrive shows how multimodal driving interactions improve short-horizon off-road pose prediction, while leaving the conversion of those gains into autonomous navigation unresolved.

Read illustrated report →
BEVDet-Tiny pipeline with multi-camera features, a depth-classification branch, outer product, calibrated point-cloud lifting, vertical pooling, a BEV encoder and a 3D detection head.
READING 498 · 6 ORIGINAL VISUALS

BEVDet

BEVDet makes camera-to-BEV detection effective through separate augmentation in each view space, then improves small-object suppression and pooling efficiency.

Read illustrated report →
BioBERT pipeline: BERT weight initialization and PubMed/PMC corpora feed biomedical pretraining, followed by named-entity, relation-extraction and question-answering fine-tuning.
READING 499 · 6 ORIGINAL VISUALS

BERT applications review

Koroteev’s review shows how a pretrained language encoder supports varied NLP uses, while its reproduced comparisons make training choices and evidence quality central to interpreting success.

Read illustrated report →
Three observed Atari frames enter blue encoders; posterior and prior categorical states connect to a purple recurrent chain, orange reward predictors, and brown image decoders. An inset depicts multiple one-hot categorical variables.
READING 500 · 5 ORIGINAL VISUALS

DreamerV2

Discrete latent predictions and balanced KL learning make a separately trained world model useful for Atari policy learning, while leaving important per-game failures and configuration ambiguities.

Read illustrated report →
Two-stage architecture: a semantic point cloud is projected into guidance for a stochastic encoder-decoder; its structural predictions and RGB guidance condition a Multi-SPADE image generator. The diagram includes a shared encoder, prior/posterior noise branches, loss blocks and a no-gradient marker.
READING 501 · 6 ORIGINAL VISUALS

Pathdreamer

Pathdreamer turns remembered geometry into stochastic future panoramas that help a separate navigation planner, but plausible room completions can still misrepresent the real route.

Read illustrated report →
A three-column table maps uncertainty, score variability and aggregate-metric weaknesses to bootstrap intervals, performance profiles and IQM.
READING 502 · 6 ORIGINAL VISUALS

The Statistical Precipice

Reliable few-run RL comparisons require matched evaluation protocols, uncertainty intervals and complementary summaries of the run-score distribution.

Read illustrated report →
Two encoders map paired images and texts to a similarity matrix with diagonal positives; class descriptions then become text embeddings used to classify a new image.
READING 503 · 6 ORIGINAL VISUALS

CLIP

Learning which image and text belong together creates a language-defined visual classifier, but broad transfer still depends on data coverage, label design and evaluation protocol.

Read illustrated report →
Two rows of differently shaped cabinets, swivel chairs and buckets illustrate geometric and articulation diversity.
READING 504 · 5 ORIGINAL VISUALS

ManiSkill

ManiSkill makes object-level manipulation generalization measurable with curated simulation and successful demonstrations, while its baselines reveal how little fixed-object competence guarantees transfer.

Read illustrated report →
Three road sketches: alternative left/right turns at an intersection, early and late merging paths, and alternative overtaking paths amid other traffic. The observed ego route is white and the hypothetical planner route is red; panels are labeled a, b and c.
READING 505 · 2 ORIGINAL VISUALS

nuPlan

nuPlan proposes evaluating a planner through the driving consequences of its trajectories, with the usefulness of that evaluation depending on agent simulation and an unfinished scoring specification.

Read illustrated report →
Three panels show horizontal node-displacement and trajectory-action histograms, with mesh holes and obstructing furniture between them.
READING 506 · 6 ORIGINAL VISUALS

VLN-CE

Removing the navigation graph exposes long-horizon control and observation failures that instruction-following scores can otherwise hide.

Read illustrated report →
A circular learning loop and three horizontal panels connect real-game policy interaction, world-model training and policy training inside the model.
READING 507 · 6 ORIGINAL VISUALS

SimPLe

SimPLe turns limited Atari experience into a learned simulator for policy training, exchanging real interactions for computation while controlling prediction drift through stochastic latents and short rollouts.

Read illustrated report →
Two panorama graphs grouped into four rooms; the right panel highlights the route p8, p13, p10, p7, p11, p6 through rooms r0, r2 and r3.
READING 508 · 6 ORIGINAL VISUALS

Room-Across-Room (RxR)

RxR makes multilingual route following a richer grounding problem by pairing less predictable paths with synchronized human demonstrations, while its baseline agents remain far below human performance.

Read illustrated report →
Top view of the collection car with six green camera labels, five blue radar labels, a roof lidar, an IMU and local coordinate-axis markers.
READING 509 · 6 ORIGINAL VISUALS

nuScenes

nuScenes makes sensor coverage, temporal context and evaluation rules part of the perception benchmark, so a detector’s ranking must be read together with its inputs and error profile.

Read illustrated report →
Three rows of eight simulated manipulation scenes, including block stacking, appliance interaction, a checkers board and plant watering.
READING 510 · 6 ORIGINAL VISUALS

RLBench

RLBench makes manipulation tasks, observations and generated demonstrations reusable, while leaving learner design and the proof of generalization to subsequent experiments.

Read illustrated report →
Top-down vehicle with five LiDAR positions, an enlarged five-camera arrangement, and red x, green y and blue upward z axis markers.
READING 511 · 6 ORIGINAL VISUALS

Waymo Open Dataset (CVPR 2020)

Synchronized multisensor sequences and independent spatial labels make perception scale and geographic transfer measurable, while protocol ambiguities limit exact reproduction.

Read illustrated report →
A three-row montage of daytime and nighttime driving scenes with scene tags, boxes, lane and area overlays, full-frame masks, and sequential object masks.
READING 512 · 6 ORIGINAL VISUALS

BDD100K

BDD100K shows that plentiful simpler annotations can improve scarce complex perception tasks, but transfer depends on task pairing, data scale, and the error being measured.

Read illustrated report →
A blue original sentence has two crossed-out spans; arrows replace them with distinct sentinels in a green input, while the pink target contains the removed words and sentinels.
READING 513 · 6 ORIGINAL VISUALS

T5

T5 makes diverse language tasks comparable through text-to-text generation, then improves transfer with economical denoising, broad data, and scale at substantial training and deployment cost.

Read illustrated report →
Four-by-four montage of sixteen game screenshots, including platforms, spacecraft, corridors, mazes, underwater scenes and moving obstacles.
READING 514 · 6 ORIGINAL VISUALS

Procgen Benchmark

Procedurally varied game levels expose memorization that training scores can hide, while wider visual policies improve generalization at a higher, unquantified compute cost.

Read illustrated report →
Three rows compare ML1 goal variation, MT10 repeated task families, and ML10 training families versus five different test families, separated by train and test columns.
READING 515 · 6 ORIGINAL VISUALS

Meta-World

Meta-World makes generalization measurable by combining shared manipulation structure with distinct task families, exposing the gap between learning known skills and adapting to new ones.

Read illustrated report →
Three blue prompt panels show zero, one and several English–French examples; the comparison column inserts yellow gradient updates between training examples.
READING 516 · 6 ORIGINAL VISUALS

GPT-3

Scaling text pretraining makes a fixed language model better at following demonstrations in its context, with substantial gains but uneven task coverage and costly inference.

Read illustrated report →
Three panels show observation-conditioned dynamics learning, latent trajectories with predicted actions/rewards/values, and observation-conditioned action execution.
READING 517 · 6 ORIGINAL VISUALS

Dreamer

Dreamer turns short latent rollouts into a policy-learning signal through bootstrapped values, gaining efficient visual control while remaining dependent on the learned dynamics and representation.

Read illustrated report →
A red initial cue leads through a chain of grey states to a choice between red and blue endpoints.
READING 518 · 6 ORIGINAL VISUALS

bsuite

bsuite makes RL capabilities testable through fixed, scalable behavioural experiments, but its scores remain local diagnostics and the paper contains unresolved protocol discrepancies.

Read illustrated report →
Three columns show a human selecting summary j over k, a reward model learning from their score difference, and a policy receiving a summary-level reward for PPO updates.
READING 519 · 6 ORIGINAL VISUALS

Summarization from Human Feedback

Learning a reward from human comparisons improves summary generation, but optimizing that reward too aggressively can reverse the improvement.

Read illustrated report →
Trajectory montages labeled Sawyer, Franka, WidowX, Kuka, Baxter, Google Robot and Fetch, with arrows indicating time from left to right.
READING 520 · 6 ORIGINAL VISUALS

RoboNet

Shared robot experience can reduce target-robot data needs, but the tested models sometimes benefit more from a relevant subset than from the full diversity of RoboNet.

Read illustrated report →
Upstream data trains a model inside algorithm A; a shared adaptation procedure branches to separate tasks sampled from P_T, each with adaptation images, test images and an evaluation score.
READING 521 · 5 ORIGINAL VISUALS

VTAB

VTAB measures representation quality through adaptation to diverse unseen tasks, exposing how upstream supervision, tuning budget and the transfer strategy jointly determine apparent progress.

Read illustrated report →
Two pairs of black reference and query paths with endpoint labels and dashed links showing optimal ordered DTW correspondences.
READING 522 · 5 ORIGINAL VISUALS

nDTW and SDTW

Ordered, distance-aware trajectory alignment makes route fidelity measurable, while success gating and reward design determine how that fidelity affects navigation evaluation.

Read illustrated report →
Original ten-line MBPO pseudocode: fit a model on environment data, collect a policy action, branch short rollouts from replay states, and update the policy on model data.
READING 523 · 5 ORIGINAL VISUALS

MBPO

MBPO gains useful training data from short model rollouts anchored to real experience, trading longer imagination for more reliable local predictions.

Read illustrated report →
A street image with sparse visual marks accompanies a LiDAR point display: red edge points, blue plane points and a green estimated trajectory.
READING 524 · 5 ORIGINAL VISUALS

LIC-Fusion

LIC-Fusion combines sparse visual and LiDAR constraints with inertial propagation and online calibration, improving reported odometry accuracy while leaving the contribution of each component unisolated.

Read illustrated report →
Three horizontal bands show datasets, simulators and tasks; upward arrows inside Habitat connect Generic Dataset Support to Habitat Sim and Habitat API.
READING 525 · 6 ORIGINAL VISUALS

Habitat

A modular, fast simulator makes training scale and dataset choice experimentally visible, while the navigation conclusions remain conditional on idealized sensing and a specific benchmark protocol.

Read illustrated report →
Camera views surround an aligned LiDAR scene with green vehicle cuboids, magenta lane centerlines, orange driveable-region markings and a small map inset.
READING 526 · 6 ORIGINAL VISUALS

Argoverse

Argoverse makes road geometry an explicit prior for tracking and forecasting, while its baseline protocols leave some map benefits entangled with candidate generation and metric choice.

Read illustrated report →
A fixed skill prior feeds a learned policy and discriminator; environment actions generate next states. Adjacent pseudocode specifies skill sampling, intrinsic reward and alternating updates.
READING 527 · 6 ORIGINAL VISUALS

DIAYN

DIAYN turns state–skill discriminability into an intrinsic reward, building reusable stochastic policies whose downstream usefulness still depends on what the discriminator can distinguish.

Read illustrated report →
A narrated furniture-finishing clip feeds a green video network and a blue text network; both arrows terminate at nearby points in a joint embedding.
READING 528 · 6 ORIGINAL VISUALS

HowTo100M

Weak subtitle alignment at web-video scale can train a useful text–video embedding, but its transfer depends on negative sampling, domain and target adaptation.

Read illustrated report →
Three temporal graphical models compare a deterministic RNN, stochastic SSM, and RSSM with both square memory nodes and circular latent-state nodes.
READING 529 · 6 ORIGINAL VISUALS

PlaNet

PlaNet turns pixel histories into a compact stochastic dynamical model and uses online action search to gain data efficiency, while retaining finite-horizon planning costs and task-dependent performance.

Read illustrated report →
Four rows of StarCraft 2 frames show a moving unit, mineral collection, an army battle and medivac transport; time advances left to right.
READING 530 · 6 ORIGINAL VISUALS

FVD and StarCraft 2 Videos

FVD uses action-recognition features to compare whole-video distributions, gaining sensitivity to temporal defects while retaining dependence on representation, sampling and protocol.

Read illustrated report →
Original Equation (1): one over N times a sum over episodes i of success indicator S_i multiplied by shortest-path length ℓ_i divided by max(p_i, ℓ_i).
READING 531 · 1 ORIGINAL VISUALS

SPL and Navigation Evaluation

SPL rewards explicitly completed navigation and efficient paths, but meaningful comparison also requires matching goals, sensors, exploration access and success rules.

Read illustrated report →
Environment observations enter VAE V; latent z and recurrent state h enter controller C; actions return to both the environment and MDN-RNN M.
READING 532 · 6 ORIGINAL VISUALS

World Models

A separately learned visual and recurrent world model makes a small controller effective, but training inside its predictions requires guarding against exploitable simulator errors.

Read illustrated report →
Two plots of tolerance versus x, each flat at one between zero and one, with different tails outside the target bounds.
READING 533 · 5 ORIGINAL VISUALS

DeepMind Control Suite

A standardized physics benchmark makes continuous-control scores interpretable, while its baselines show why observation modality and training budget must remain part of every comparison.

Read illustrated report →
Two sequence diagrams: VLN encodes an instruction and outputs actions while receiving successive room images; VQA encodes a question and outputs answer words from an image.
READING 534 · 5 ORIGINAL VISUALS

R2R: language-guided navigation

R2R couples human route instructions to a deterministic simulator built from real panoramas, revealing that learning to act in familiar buildings transfers poorly to unseen ones.

Read illustrated report →
Two patches pass through feature network F, then normalization, subtraction, channel weighting and aggregation yield d_0. A separate branch maps d_0 and d_1 through G to a cross-entropy judgment loss.
READING 535 · 5 ORIGINAL VISUALS

LPIPS

Deep feature distances track local human similarity judgments, but improving their fit to synthetic distortions can weaken transfer relative to simple calibration.

Read illustrated report →
Two network diagrams: an actor above a critic, each combining a recurrent state/action branch with a feedforward branch. Simulator parameters enter only the critic.
READING 536 · 6 ORIGINAL VISUALS

Dynamics Randomization

A recurrent policy trained across randomized simulated physics transfers puck pushing to a real robot, but the evidence remains tied to state-based sensing and a small physical evaluation.

Read illustrated report →
Table comparing dataset year, action classes, clips per class, total clips and source-video counts for HMDB-51, UCF-101, ActivityNet-200 and Kinetics.
READING 537 · 6 ORIGINAL VISUALS

Kinetics: building an action benchmark

Kinetics combines large-scale clip collection with human verification and model-assisted cleanup, exposing complementary appearance and motion signals while retaining label ambiguity and selection-bias questions.

Read illustrated report →
Four views of the same reconstructed building show its texture, separate object instances, raw object categories and canonical 40-category labels.
READING 538 · 6 ORIGINAL VISUALS

Matterport3D

Globally aligned panoramic scans turn whole buildings into reusable perception supervision, but their value depends on viewpoint coverage, annotation choices and the quality of evaluation targets.

Read illustrated report →
Three pose circles x1, x2 and x3 form a left-to-right motion chain. Dashed bearing links connect x1 and x2 to landmark l1, and x3 to landmark l2.
READING 539 · 5 ORIGINAL VISUALS

Factor Graphs for Robot Perception

Conditioned sensor models turn SLAM into a sparse estimation problem, but adequate constraints and trustworthy local optimization remain essential.

Read illustrated report →
An image passes through a CNN encoder, a nearest-neighbor dictionary lookup and a CNN decoder. A red backward arrow bypasses the discrete index grid; a separate embedding diagram shows the selected e2 and encoder output.
READING 540 · 6 ORIGINAL VISUALS

VQ-VAE

VQ-VAE makes discrete representations trainable by separating nearest-neighbor encoding, gradient routing and prior learning, trading exact reconstruction for a compact space that can support structured generation.

Read illustrated report →
Four rows of six video frames: placing a remote in a cardboard box, pretending to place candy on a chair, pushing a green chilli off a table, and moving a puncher closer to scissors. Each row has its completed description, with object phrases in italics.
READING 541 · 5 ORIGINAL VISUALS

Something-Something (ICCV 2017)

Recording fine-grained actions from structured language prompts creates a demanding video-recognition dataset, but recognition errors alone do not establish physical common sense.

Read illustrated report →
A ball moves down an objective landscape, passes a narrow minimum marked theta-plus, and settles around the broad minimum theta-star.
READING 542 · 6 ORIGINAL VISUALS

TTUR and FID

Separate GAN learning rates can improve selected training outcomes, while FID measures image-distribution similarity through feature moments; both claims require careful attention to their assumptions and evaluation protocol.

Read illustrated report →
PointNet diagram with a blue classification path, two matrix-predicting T-nets, max pooling into a global feature, and a yellow segmentation path receiving both point features and global context.
READING 543 · 6 ORIGINAL VISUALS

PointNet

Shared pointwise features and max pooling turn unordered 3D points into useful recognition representations, with learned alignment and a conditional critical-point account of robustness.

Read illustrated report →
Four third-person views of one Town 2 street: clear daylight, rain, wet pavement after rain, and sunset.
READING 544 · 6 ORIGINAL VISUALS

CARLA

CARLA makes urban-driving experiments controllable and diagnosable, while its original baselines expose a large gap between reaching destinations in familiar scenery and driving reliably in a new town.

Read illustrated report →
Eight MuJoCo reward-versus-timestep panels compare direct true-reward RL in orange, three synthetic-query budgets in blue shades, and human feedback in purple; the legend says 750 human queries.
READING 545 · 4 ORIGINAL VISUALS

Deep RL from Human Preferences

A separately learned reward model turns sparse comparisons into a dense RL training signal, but its usefulness depends on continuing feedback and on what humans can infer from short clips.

Read illustrated report →
An eagle photograph, a list of reference descriptions R1 through R50, and two candidate captions connected conceptually by a triplet annotation box assigning a reference to A and the candidates to B and C.
READING 546 · 5 ORIGINAL VISUALS

CIDEr

CIDEr measures caption consensus by averaging informative n-gram matches across human references, gaining reliability from denser annotations while inheriting their definition of a typical description.

Read illustrated report →
A shaded observed x node and unshaded latent z node sit in an N plate; solid arrows run from z to x and from theta to both, while dashed arrows lead from x and phi to z.
READING 547 · 5 ORIGINAL VISUALS

Auto-Encoding Variational Bayes

AEVB trains recognition and generation through a differentiable latent-sampling transformation, gaining efficient approximate inference while retaining a variational gap and a restricted experimental posterior family.

Read illustrated report →
Freeway in eight-colour SECAM format beside a coarse grid showing BASS colour occupancy after background removal.
READING 548 · 6 ORIGINAL VISUALS

Arcade Learning Environment

ALE turns varied Atari games into a shared evaluation protocol, showing that representation, simulator access, and score aggregation can each change which agent appears strongest.

Read illustrated report →
Four photographs show the fr1 office, fr2 industrial hall, a handheld Kinect with reflective markers, and a Pioneer robot carrying a marked Kinect.
READING 549 · 6 ORIGINAL VISUALS

RGB-D SLAM Benchmark

Independent camera-pose measurements make RGB-D trajectory comparisons reproducible, provided calibration, temporal association and the chosen error metric remain explicit.

Read illustrated report →
A sensor-equipped car, a recorded trajectory, a road image with disparity and flow maps, and projected 3D boxes around parked vehicles.
READING 550 · 6 ORIGINAL VISUALS

KITTI (CVPR 2012)

KITTI exposes road-scene perception failures through calibrated references and diagnostic evaluation, while its selected scenes and partial labels limit claims about autonomous driving readiness.

Read illustrated report →
Three stages of a history tree: simulated action and observation branches, the realized a2/o2 path, and the retained subtree with particle sets.
READING 551 · 4 ORIGINAL VISUALS

POMCP

POMCP shares simulator trajectories between history-tree planning and particle-belief updates, gaining scalability while relying on finite samples and a supplied generative model.

Read illustrated report →
Table of Chinese system-level Pearson correlations: BLEU 0.817, NIST 0.892, precision 0.752, recall 0.941, F1 0.948, Fmean 0.952, and METEOR 0.964.
READING 552 · 6 ORIGINAL VISUALS

METEOR (2005)

Explicit lexical alignment and recall-weighted coverage improve agreement with human translation judgments, while sentence-level agreement remains modest and sensitive to how those judgments are processed.

Read illustrated report →
Five corridor panels show three doors and a moving robot above location-belief curves. A flat belief becomes three peaks after sensing, shifts and spreads after motion, concentrates near the middle door after another observation, and spreads again during further travel. Red observation-likelihood curves accompany the sensing panels.
READING 553 · 2 ORIGINAL VISUALS

Probabilistic Robotics — Chapter 1

Representing uncertain location as a distribution lets a robot combine sensing with motion and choose informative routes, at the cost of computation and approximation.

Read illustrated report →
Six Boat panels: original, contrast stretched and mean shifted across the top; JPEG blocks, blur and salt-and-pepper noise across the bottom.
READING 554 · 6 ORIGINAL VISUALS

SSIM

SSIM compares local brightness, contrast and normalized structure to better track perceived compression quality, while retaining the need for an aligned reference and application-specific validation.

Read illustrated report →
Reference X contains A B C D E F G. Candidate Y1 places matching A B C D consecutively; Y2 separates those same matches with H, K and I. Underlines identify the matching tokens.
READING 555 · 5 ORIGINAL VISUALS

ROUGE

ROUGE turns reference-summary overlap into inexpensive evaluation scores, but its agreement with human content judgments depends on the matching rule, reference policy and summarization task.

Read illustrated report →
An inverted pendulum on a cart appears beside a biped on uneven ground. The biped COG is marked x_G; a dashed line connects it to x_ZMP on a horizontal support line.
READING 556 · 6 ORIGINAL VISUALS

Humanoid Control through ZMP Manipulation

An inverted-pendulum approximation turns feasible ZMP references into whole-body joint motion, trading explicit full-body dynamics for a compact feedback pipeline demonstrated only in simulation.

Read illustrated report →
Grouped precision bars for n-gram orders one through four, with H2, H1, S3, S2 and S1 identified in the legend.
READING 557 · 6 ORIGINAL VISUALS

BLEU

BLEU makes translation comparison repeatable by pooling clipped phrase matches and penalizing corpus brevity, at the cost of dependence on reference wording and evaluation protocol.

Read illustrated report →
The world sends an observation to state estimator SE; belief b enters policy pi, whose action affects the world and feeds back to SE.
READING 558 · 6 ORIGINAL VISUALS

POMDP planning and acting

A probability distribution over hidden states makes principled planning possible, but exact contingent policies can be costly to compute and may require substantial memory.

Read illustrated report →