JailWAM: Jailbreaking World Action Models in Robot Control
Classification
View four quadrants- Major category
- Benchmarks & simulators
- Quadrant
- Not applicable
- Architecture
- Not applicable
- Prediction paradigm
- Not applicable
- Source review status
- Not assigned
Category review. The paper evaluates existing WAM policies using instruction attacks, trajectory rendering, learned risk screening and human-verified closed-loop simulated outcomes. It is an evaluation framework rather than a new action policy. The cataloged contribution is a dataset/data-generation method or evaluation/simulation resource, not the architecture of an evaluated or external policy. Reading evidence
Contribution
JailWAM evaluates instruction-induced robot risk by rendering predicted actions as trajectory charts, screening them with a trained vision-language discriminator, and verifying selected candidates in closed-loop simulation. It reports 84.20% human-verified attack success on LingBot-VA, but its cheaper screening pipeline misses some unsafe executions. Success includes motion failure as well as catastrophic risk.
Abstract
An abstract has not been added yet.
Affiliations
MoE Key Lab of Artificial Intelligence, AI Institute, Shanghai Jiao Tong University, Shanghai, China; PR Lab, Nanjing University, Suzhou, China; Defense Innovation Institute, Chinese Academy of Military Science, Beijing, China