RESEARCH PAPER

CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models

Qingqing Zhao; Yao Lu; Moo Jin Kim; Zipeng Fu; Zhuoyang Zhang; Yecheng Wu; Zhaoshuo Li; Qianli Ma; Song Han; Chelsea Finn; Ankur Handa; Ming-Yu Liu; Donglai Xiang; Gordon Wetzstein; Tsung-Yi Lin

Classification

View four quadrants
Major category
WAMs
Architecture
One Model
Prediction paradigm
IDM
Source review status
Not assigned

Category review. A shared multimodal model explicitly generates a task-conditioned future subgoal image before predicting and executing the action chunk conditioned on that image. The shared multimodal model first generates a subgoal image, then predicts an action chunk conditioned on it. Reading evidence

AT A GLANCE

Contribution

CoT-VLA makes a 7B multimodal model generate a future subgoal image before predicting a robot action chunk. The shared model learns image prediction from demonstrations and captioned videos, while action supervision requires demonstrations. Experiments support useful goal conditioning and stronger average manipulation performance, with uneven task gains, substantial inference overhead and reporting inconsistencies. Visual chain of thought here means an explicit predicted image used for control; general reasoning is not independently established.

Abstract

An abstract has not been added yet.

Affiliations

NVIDIA; Stanford University; MIT