CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models
Classification
View four quadrants- Major category
- WAMs
- Quadrant
- Q2 · One Model × IDM
- Architecture
- One Model
- Prediction paradigm
- IDM
- Subcategories
- Visual planning & IDM
- Source review status
- Not assigned
Category review. A shared multimodal model explicitly generates a task-conditioned future subgoal image before predicting and executing the action chunk conditioned on that image. The shared multimodal model first generates a subgoal image, then predicts an action chunk conditioned on it. Reading evidence
Contribution
CoT-VLA makes a 7B multimodal model generate a future subgoal image before predicting a robot action chunk. The shared model learns image prediction from demonstrations and captioned videos, while action supervision requires demonstrations. Experiments support useful goal conditioning and stronger average manipulation performance, with uneven task gains, substantial inference overhead and reporting inconsistencies. Visual chain of thought here means an explicit predicted image used for control; general reasoning is not independently established.
Abstract
An abstract has not been added yet.
Affiliations
NVIDIA; Stanford University; MIT