RESEARCH PAPER

DUET-DINO: Simultaneous Cross-View World Modeling for Latent Planning in Robot Manipulation

Nisarga Nilavadi; Ralf Römer; Moritz Reuss; Michael Krawez; Tobias Jülg; Angela P. Schoellig; Rudolf Lioutikov; Wolfram Burgard

Classification

View four quadrants
Major category
WAMs
Architecture
Pending verification
Prediction paradigm
Other mechanisms
Source review status
Not assigned

Category review. Cross-view latent dynamics predict candidate-action consequences; CEM searches and executes robot controls against paired image goals. Its use of frozen DINO features does not make the complete planning method a visual encoder component. The described action mechanism is external planning, model-assisted policy optimization, geometric tracking, or video-conditioned control; the source does not establish joint future/action generation or an IDM action decoder. Reading evidence

AT A GLANCE

Contribution

DUET-DINO couples side- and wrist-camera latent predictors through cross-attention, then searches seven-dimensional robot actions against paired goal images. Its clearest evidence is improved simulated spatial and orientation planning; hardware success remains limited under a smaller planning budget and safety terminations (e-architecture, e-planning, e-angled, e-hardware).

Abstract

An abstract has not been added yet.

Affiliations

Artificial Intelligence and Robotics Lab, University of Technology Nuremberg (UTN), Germany; Learning Systems and Robotics Lab, Technical University of Munich (TUM), Germany; Intuitive Robots Lab, Karlsruhe Institute of Technology (KIT), Germany; NVIDIA; Robotics Institute Germany