RESEARCH PAPER

JailWAM: Jailbreaking World Action Models in Robot Control

Hanqing Liu; Songping Wang; Jiahuan Long; Jiacheng Hou; Jialiang Sun; Chao Li; Yang Yang; Wei Peng; Xu Liu; Tingsong Jiang; Yao Mu; Wen Yao

Classification

View four quadrants
Architecture
Not applicable
Prediction paradigm
Not applicable
Source review status
Not assigned

Category review. The paper evaluates existing WAM policies using instruction attacks, trajectory rendering, learned risk screening and human-verified closed-loop simulated outcomes. It is an evaluation framework rather than a new action policy. The cataloged contribution is a dataset/data-generation method or evaluation/simulation resource, not the architecture of an evaluated or external policy. Reading evidence

AT A GLANCE

Contribution

JailWAM evaluates instruction-induced robot risk by rendering predicted actions as trajectory charts, screening them with a trained vision-language discriminator, and verifying selected candidates in closed-loop simulation. It reports 84.20% human-verified attack success on LingBot-VA, but its cheaper screening pipeline misses some unsafe executions. Success includes motion failure as well as catastrophic risk.

Abstract

An abstract has not been added yet.

Affiliations

MoE Key Lab of Artificial Intelligence, AI Institute, Shanghai Jiao Tong University, Shanghai, China; PR Lab, Nanjing University, Suzhou, China; Defense Innovation Institute, Chinese Academy of Military Science, Beijing, China