RESEARCH PAPER

This&That: Language-Gesture Controlled Video Generation for Robot Planning

Boyang Wang; Nikhil Sridhar; Chao Feng; Mark Van der Merwe; Adam Fishman; Nima Fazeli; Jeong Joon Park

Classification

View four quadrants
Major category
WAMs
Architecture
Pending verification
Prediction paradigm
Other mechanisms
Source review status
Not assigned

Category review. Language and gestures condition an explicit future-video plan, and a separate DiVA policy maps that plan plus live observations into executable action chunks. The described action mechanism is external planning, model-assisted policy optimization, geometric tracking, or video-conditioned control; the source does not establish joint future/action generation or an IDM action decoder. Reading evidence

AT A GLANCE

Contribution

This&That turns an initial image, language and pointing coordinates into a generated video plan, then uses a separate DiVA behavior-cloning controller to follow it with live feedback. Gestures improve intent disambiguation in Bridge video generation and simulated block manipulation; real-robot execution remains untested.

Abstract

An abstract has not been added yet.

Affiliations

University of Michigan; University of Washington