Attacking the Trusted Imagination: Oracle-Level Integrity Attacks on Imagine-then-Act World Models
1. Paper overview
In one sentence: Observation attacks can corrupt imagined futures and impair a simulated MPC, but latent damage, detector evasion and imagination-specific task failure require separate evidence. e-threate-corruptione-targetede-mpce-causal
| At a glance | What to know |
|---|---|
| Research problem | Source description A policy can tolerate corrupted imagined features while a planner or verifier makes decisions from those predictions. The paper asks whether white-box observation perturbations compromise this trusted intermediate representation, and whether latent damage translates into independent decisions or executed-task failure. e-threat |
| Core mechanism | Source description The work supplies an observation-to-imagination attack, corruption/steering comparison, clean-reference-free denoiser detector, and controls separating latent fidelity from task success. e-pipelinee-objectivese-detectione-reactive |
| A key reported result | LaDi-WM MPC on LIBERO-10 task 4: Adversarial 1/20 = 0.05; Wilson 95% interval [0.01, 0.24]. Closed-loop success rate. Simulation; 20 episodes per condition; epsilon 0.01. Random 14/20 = 0.70 [0.48, 0.85]; clean 0.55; reported Fisher p < 10^-4. An attack-versus-random separation on one task. At epsilon 0.03 rates are 0.05 versus 0.65; at 0.06, 0 versus 0.40. Larger noise also harms the baseline. e-protocole-mpc |
| Reading caution | Reader analysis Evaluation is simulation only; MPC failure is restricted to task 4. The Figure 1 contrast combines different architectures and budgets, so it is not a controlled estimate of the causal effect of using imagination. e-limitse-threate-reactivee-mpc |
Core contributions
- Source description
The work supplies an observation-to-imagination attack, corruption/steering comparison, clean-reference-free denoiser detector, and controls separating latent fidelity from task success. e-pipelinee-objectivese-detectione-reactive
- Author claim
The authors explain the asymmetry through a natural-future manifold and argue that adaptive evasion requires sacrificing corruption. This is a mechanism interpretation supported by limited probes, not a robustness theorem. e-mechanisme-adaptive
Figure 2. The perturbation changes the predicted future through frozen model components. Original paper, p. 4 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. Follow the solid arrows from the perturbed observation through E and W to the imagined latent. The red dashed arrow returns a loss gradient to the observation perturbation; it does not update model weights. The yellow box names the two objectives: move away from the clean imagination or approach a specified target. Read the detector's comparison symbol with Algorithm 1 on page 5: corruption is flagged when its score is greater than the threshold. The algorithm subtracts the signed gradient; untargeted divergence increases because its loss has a minus sign. The continuous, cache-safe path is a condition for this gradient attack, not a property of every target in Table 2. e-pipelinee-objectivese-algorithme-threate-models
What it supports. The attack exploits access to a differentiable prediction pathway without requiring access to the downstream oracle's internals. A separate denoiser score can inspect the resulting latent without a clean reference. The diagram therefore distinguishes where the perturbation is optimized from where a later decision may be affected.
Where the evidence stops. The oracle and detector are drawn as parallel branches; the drawing alone does not specify an enforced blocking policy. Its no-VQ condition excludes the discrete RynnVLA imagination path from this direct continuous formulation.
2. Motivation
2.1 The problem and the proposed response
A policy can tolerate corrupted imagined features while a planner or verifier makes decisions from those predictions. The paper asks whether white-box observation perturbations compromise this trusted intermediate representation, and whether latent damage translates into independent decisions or executed-task failure. e-threat
2.2 What this reading follows
Imagine a robot whose world model forecasts what happens next. A reactive controller may recover from a poor forecast through environmental feedback, while a verifier or planner can make a decision from the forecast before any correction is available. This paper attacks that trusted prediction by changing the current camera observation. Its strongest contrast is between easy visible corruption and limited cross-scene steering. Its strongest rollout result is a large attack-versus-random gap on one LaDi-WM task. Read the six visuals as an evidence chain: how the attack enters, what the forecast becomes, how detection responds, and which causal conclusions the task controls permit. e-threate-corruptione-targetede-mpce-causal
3. Research context
We place the paper in the collection through its world–action interface. The catalog labels and the reading’s assessment are shown separately.
| Catalog dimension | Recorded classification |
|---|---|
| Major category | Foundational work |
| Architecture | Not applicable |
| Prediction paradigm | Not applicable |
| Quadrant | Not applicable |
This table preserves the labels recorded at reading time. The current major category is Related resources. View the current classification.
3.1 Evidence-based assessment
Supports the recorded classification
Foundational work / Evaluation metrics & protocols fits an attack-and-defense evaluation of existing WAMs. This paper introduces no new WAM architecture to place in a prediction quadrant. A shared backbone alone, illustrated by RynnVLA, does not imply that actions consume predicted futures; the recorded architecture and quadrant remain Not applicable. e-threate-modelse-rynn
This is the collection’s architectural analysis, not a new related-work survey. Benchmark comparisons and their protocols appear in Section 6.
4. Problem formulation
4.1 Inputs and outputs
| Inputs | Outputs |
|---|---|
|
|
4.2 Equations and their role
5. Method in detail
5.1 First separate the prediction from the action that consumes it
Start with the consumer, because a generated future is not automatically an action plan. RynnVLA-002 shares a backbone between image and action generation, yet the action head reads its own token state. LingBot-VA has a more direct dependency: action tokens attend the imagined latent through block-causal attention. LaDi-WM uses predicted feature evolution to refine diffusion-policy actions, and the evaluated wrapper scores candidate actions through imagination. The attack freezes these models and changes the observation, so its optimization is neither WAM training nor a new action-learning objective. Reader interpretation: this architectural range makes the negative controls informative, but also prevents Figure 1's cross-model contrast from isolating a single design choice. To reason about control, trace the dependency actually used at inference and then ask what environmental feedback can correct. e-modelse-algorithme-rynne-mpce-threat
5.2 Then distinguish leaving a clean prediction from reaching a desired one
The untargeted objective only rewards disagreement with the clean imagined latent. The targeted objective must reduce distance to a particular alternate imagination, which may depict a different scene. The authors interpret this as an off-manifold versus on-manifold distinction. Figures 3 and 4 provide a useful check on that interpretation: modest target-distance closure leaves the original scene recognizable, whereas untargeted optimization visibly damages decoded objects. The denoiser asks whether the attacked future is self-consistent under the learned prior without consulting a clean reference. Reader interpretation: Table 7 supports a tradeoff for the tested penalty sweep, but the manifold explanation is stronger than the measured cosine and score changes alone. A harmful, plausible alternate future could be a different failure mode; its exclusion would require an independent semantic or task-level test. e-objectivese-targetede-corruptione-detectione-adaptivee-mechanism
5.3 Finally carry the right uncertainty into the task-level claim
Table 8 moves beyond predicted pixels by counting completed simulated episodes. At the smallest budget, fourteen random-noise successes versus one adversarial success is substantial evidence of an optimized-direction effect on the evaluated task. However, Table 9 asks a different question at a larger budget: where does the corrupted observation cause harm? Recovery with clean execution and attacked scoring localizes the effect away from the outer ranking stage. It does not separate direct vision from imagination refinement inside execution. Reader interpretation: retain both the positive task result and this attribution limit. Likewise, do not substitute Table 6's perfect AUC at selected budgets for the full-scale safety-gate AUC of 0.771. Those are different evaluations. A robust reading follows each metric to its consumer, sample size and intervention before generalizing. e-mpce-causale-detectione-limitse-protocol
5.4 Training and inference
During training
The attack freezes the transformer, text encoder and VAE; it optimizes the input rather than retraining the WAM. Reset calls require no_grad to avoid reused text-context computation graphs. SPSA is a fallback when in-place cache operations break a required gradient. e-algorithm
The universal perturbation is optimized on tasks [0, 3], then frozen for held-out tasks [4, 7]. The detector adds no learned parameters; underlying model pretraining recipes are not provided here. e-transfere-detectione-models
During inference
After attack optimization, the consumer uses the altered imagined future. The LaDi-WM sampling MPC considers K = 4 candidate actions every five steps and executes the highest-scoring choice. These are simulated rollouts; the safety-gate and verifier evidence has a narrower decision-probe scope. e-mpce-threate-limits
5.5 Implementation flow
- Identify the actual consumer
RynnVLA-002 actions read their own token state, sharing only a backbone with imagination. LingBot-VA action tokens attend generated video latents through a block-causal mask. LaDi-WM predicts DINO/SigLIP latent evolution to refine diffusion-policy actions. e-models
- Optimize the observation
Pass the perturbed observation through E and W under c. Minimize negative divergence for untargeted damage or squared target distance for steering. Project sign-gradient updates into the infinity-norm budget; the attacker neither changes model weights nor accesses oracle internals. e-threate-pipelinee-objectivese-algorithm
- Check the predicted future
Score mean future-frame denoiser velocity-prediction norm at source noise levels t = 0.8, 0.9, 0.97. Flag scores above tau. This detector needs no clean future, unlike the attack's clean-reference fidelity metric. e-algorithme-detectione-verifiers
6. Experiments & results
This integrity-attack study perturbs observations to damage imagined futures in existing world-action models. Corruption is easier than cross-scene steering, and a denoiser detects many tested attacks. LaDi-WM MPC suffers a substantial single-task simulation success drop, but the shared visual pathway prevents attributing that failure exclusively to imagination.
6.1 Read the original evidence
Figure 3. Movement in a latent metric does not make the attacked future depict the target scene. Original paper, p. 7 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. Read the panels from left to right: the clean imagination, the targeted attack's imagination, and the requested target. The caption identifies decoded frame 12 for task 0 to task 3 and describes agent and wrist camera views within each panel. Compare the basket, cans and overall tabletop arrangement before looking at individual artifacts. The middle panel retains the original scene, while the target has a different kitchen layout. The approximate-equality and inequality markers agree with the caption and the Table 3 decoder/verifier column. That table gives an eight-task cross-scene summary; this figure illustrates one example from the setting rather than displaying all eight outcomes. e-targetede-objectives
What it supports. At epsilon 0.2, cross-scene steering reports gap_closed 0.207 ± 0.056 versus random −0.002 ± 0.008, yet the verifier flips in 0/8 cases. The visual supports the narrower conclusion that this illustrated attempt preserves scene identity despite measurable latent movement.
Where the evidence stops. The caption calls the score a latent-cosine gap, but Eq. (2) defines gap_closed using normalized norm distance. The report follows that equation. One decoded frame also cannot establish preservation of the whole predicted trajectory.
Figure 4. Untargeted optimization produces visible damage without specifying a replacement scene. Original paper, p. 8 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. Compare corresponding tabletop objects in the clean and attacked panels. The source identifies both as decoded imagined frame 12 for task 0, with epsilon 0.2 in the adversarial panel. The caption describes smearing, ghosting and a torn region on the right; the displayed change is in a model-generated future, not a photograph of executed robot behavior. Contrast this image with Figure 3: that attack seeks a particular kitchen target, while this one only increases disagreement with the model's clean imagination. Table 4 supplies the associated five-task latent results, and Table 5 separately compares sensitivity of future and current-observation latent channels. e-corruptione-negativee-verifiers
What it supports. At epsilon 0.2, Table 4 reports corruption 0.266 ± 0.03 versus 0.013 for random noise. This example makes the latent change visually interpretable. It supports an attack on prediction fidelity, while the negative channel control provides additional evidence that the future representation is especially sensitive.
Where the evidence stops. A visibly damaged prediction is not an executed-task failure. The 100% fidelity-verifier flip elsewhere uses the same cosine targeted by the attack; the independent task classifier reports only 12% untargeted misclassification.
Table 8. Small-budget attacks sharply reduce simulated MPC success compared with matched random noise. Original paper, p. 9 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. Begin with the caption: this is task 4, twenty episodes per point, with clean success 0.55. Each row changes the observation perturbation budget while the two success columns compare random and optimized perturbations. At 0.01 the counts are fourteen successes versus one; at 0.03 they are thirteen versus one. The final row shows eight random successes and no adversarial successes, so increasing perturbation magnitude also harms the random baseline. Section 5.9 describes four candidate actions scored every five steps. Its paragraph reports the confidence intervals and significance test for the first row; those uncertainty values are not printed in this crop. e-mpce-protocole-causale-limits
What it supports. At epsilon 0.01, random success 0.70 and adversarial success 0.05 have reported Wilson intervals [0.48, 0.85] and [0.01, 0.24], with Fisher p < 10^-4. This establishes an adversarial-direction effect on simulated task completion in this setting, beyond a mere change in imagined pixels.
Where the evidence stops. This is one task, not a LIBERO-10 average or physical deployment. The source labels the largest budget a saturation regime; random success there is still 0.40. Table 9 also limits attributing failure exclusively to imagination.
6.2 Results and evaluation conditions
| Task & protocol | Reported result | Comparison & interpretation |
|---|---|---|
| LaDi-WM MPC on LIBERO-10 task 4 Simulation; 20 episodes per condition; epsilon 0.01. | Adversarial 1/20 = 0.05; Wilson 95% interval [0.01, 0.24]. Closed-loop success rate | Random 14/20 = 0.70 [0.48, 0.85]; clean 0.55; reported Fisher p < 10^-4. An attack-versus-random separation on one task. At epsilon 0.03 rates are 0.05 versus 0.65; at 0.06, 0 versus 0.40. Larger noise also harms the baseline. e-protocole-mpc |
| MPC pathway isolation Task 4; epsilon 0.06; 20 episodes each. | Full attack 0; score_only 0.60; final_only 0. Success rate | Clean 0.55; matched random 0.40. Harm follows the execution-step observation. Shared direct-vision and imagination-refinement inputs prevent imagination-only attribution. e-causal |
| LingBot-VA cross-scene steering Eight tasks; epsilon 0.2. | 0.207 ± 0.056; 0/8 flips. Gap closed; verifier flips | Random −0.002 ± 0.008; same-observation instruction swapping reaches 0.46–0.52 in a separate sweep. Decoded frames remain the original scene despite latent movement. The source does not define the ± uncertainty type. e-targeted |
| LingBot-VA untargeted corruption Five tasks; epsilon 0.05 and 0.2. | 0.135 ± 0.05 at 0.05; 0.266 ± 0.03 at 0.2. 1 minus clean/adversarial latent cosine | Random 0.002 and 0.013; the first ratio is labeled approximately 60×. Decoded damage is visible. At epsilon 0.1 the channel control reports 24.5× future/observation amplification, supporting selective latent sensitivity rather than proving task harm. e-corruptione-negative |
| Fidelity versus independent verification LingBot-VA decision probes. | Fidelity flip 100% at epsilon 0.1, tau 0.9; independent V1 untargeted misclassification 12%. Flip or task-misclassification rate | Random 0% for both reported comparisons; V1 clean accuracy 1.000. The fidelity gate thresholds the optimized cosine, so its flip is partly constructed. Independent targeted false acceptance remains 0% over eight tasks at epsilon 0.2. e-verifiers |
| Denoiser detection and adaptive evasion Eight-task detection; separate four-task adaptive sweep. | AUC 0.89/1.00/1.00 at epsilon 0.1/0.2/0.4; TPR at 5% FPR 62%/100%/100%. AUC; corruption | Surface statistics near chance. Adaptive lambda 0 to 8 reduces AUC 1.00 to 0.06 and corruption 0.279 to 0.064. Detection performance depends on budget and attacker objective; residual corruption is nonzero, and adaptive downstream task harm is not reported. e-detectione-adaptive |
| Universal latent-corruption transfer Train tasks [0, 3]; held-out tasks [4, 7]; frozen perturbation. | 0.199 Held-out corruption | Random 0.039; training 0.356 versus 0.034. Transfer is demonstrated on two held-out tasks, not broad deployment or task-failure generalization. e-transfer |
| Reactive and disconnected controls LingBot-VA: ten tasks, 143–190 episodes per condition; separate RynnVLA cross-model probe. | LingBot-VA 0.966 under imagination attack at epsilon 0.8. Task success rate | Clean 0.978. RynnVLA transferred imagination attack, random and clean all 100%; direct victim-action attack 0%. These nulls limit extrapolation from latent damage. Models, budgets and protocols differ from the MPC comparison. e-reactivee-protocole-rynn |
| Safety-gate availability Full-scale N = 179, superseding an N = 3 pilot. | AUC 0.771; absolute success loss 0.18–0.39 across false-positive rates. Gate AUC; effective-success loss | The earlier pilot's AUC 1.00 and catastrophic-collapse interpretation are withdrawn. This gate result is distinct from Table 6 latent detection; per-threshold outcomes are not tabulated. e-limits |
6.3 Ablations and diagnostic examples
Read component removals and qualitative examples within their stated evaluation conditions.
Table 7. Penalizing the detector score reduces both detectability and the measured corruption. Original paper, p. 9 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. Read downward as lambda increases from 0 to 8. In the source objective, the attacker minimizes negative corruption plus lambda times detector score, so larger lambda gives evasion more weight. The corruption column measures departure from the clean latent; det_score is the denoiser statistic; AUC summarizes clean-versus-attack score ranking. The surrounding text gives a clean score of 0.369, which helps interpret attacked scores of 0.345 and 0.337 in the last two rows. Page 6 identifies this as a four-task probe. Examine the middle rows as well as the endpoints: lowering detectability is gradual, and corruption remains positive throughout the displayed sweep. e-adaptivee-protocole-algorithm
What it supports. From lambda 0 to 8, corruption falls from 0.279 to 0.064 and AUC from 1.00 to 0.06. The observations support a tradeoff under this adaptive objective. They do not show that every attack with a low detector score becomes behaviorally harmless.
Where the evidence stops. The source gives no operating-threshold task-success results for these adaptive rows, and does not specify their epsilon in Table 7 or its paragraph. AUC below 0.5 indicates reversed score ranking, not zero residual corruption.
Table 9. The intervention localizes harm to the execution input while leaving two internal pathways entangled. Original paper, p. 10 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. Use the parenthetical descriptions to identify which observation is corrupted. In score_only, MPC ranks candidates using the adversarial image but execution receives a clean image. In final_only, ranking is clean and execution receives the adversarial image. Compare these with the clean and full-attack rows: score_only recovers to twelve successes, while final_only has none. The table calls the outer planner innocent, but the paragraph immediately above qualifies what follows from that recovery. During execution, direct vision and imagination refinement both use the same corrupted image and DINO-SigLIP encoder. This intervention separates outer scoring from execution, not the two internal contributors to execution. e-causale-mpce-limits
What it supports. At epsilon 0.06, score_only success is 12/20 = 0.60 versus clean 11/20 = 0.55, while final_only and the full attack both achieve 0/20. The result argues against outer candidate ranking as the failure carrier in this probe.
Where the evidence stops. The diagnostic uses the larger budget, outside the small-budget window emphasized in Table 8. Its shared-encoder confound prevents either a pure direct-vision or an imagination-only causal claim, and clean-like scores do not prove equivalence.
7. Analysis & limitations
7.1 What the evidence leaves open
Evaluation is simulation only; MPC failure is restricted to task 4. The Figure 1 contrast combines different architectures and budgets, so it is not a controlled estimate of the causal effect of using imagination. e-limitse-threate-reactivee-mpc
The manifold account is plausible but cosine divergence does not itself prove departure from a natural-future manifold. Four adaptive weights cannot establish impossibility of harmful on-manifold evasion. e-objectivese-mechanisme-adaptive
The paper reports unreliable semantic CLIP verification at the decoded resolution. The observation threat model and later appeal to direct latent manipulation are different interventions; the latter cannot silently explain the former's task failures. e-limitse-causale-threat
7.2 Questions for discussion
- Can an attack cause independently measured downstream harm while staying below a detector threshold fixed on held-out clean data?
- Does the MPC failure persist across tasks when direct vision and imagination refinement receive independently controlled inputs?
8. Reproducibility audit
8.1 Requirements and known gaps
Reproduction requires the three model checkpoints, LIBERO-10 setup, attack/defense implementation and rollout records. The authors offer artifacts on request and report a single-A100 cluster; neither a public code link nor software versions are supplied. e-modelse-protocole-repro
Recover exact checkpoint identifiers, observation scaling/clipping, PGD T and alpha, seeds, random-noise distribution, detector calibration and MPC scoring details. The universal magnitude lacks a clear convention. These omissions impede exact replication even though the core algorithm is specified. e-algorithme-transfere-mpce-repro
8.2 Proposed reproduction checks
The following checks are proposals motivated by the paper. They have not been run as part of this reading.
Check 1: Repeat the pathway intervention inside the small-budget failure window
Reader-proposed check, not performed: once the missing implementation details are recovered, repeat clean, matched random, full attack, score_only and final_only at epsilon 0.01 and 0.03 with identical task initializations and candidate-action seeds. Add instrumentation that independently supplies clean or attacked observations to direct vision and imagination refinement during execution, keeping weights fixed. Start on task 4 and extend to the other tasks with a prespecified episode count. Imagination-only failure with clean direct vision would support specific attribution; failure confined to the direct-vision branch would weaken it. Report paired outcomes and uncertainty, not only pooled success. e-mpce-causale-limitse-repro
Check 2: Test whether adaptive evasion remains harmless at a fixed gate threshold
Reader-proposed check, not performed: calibrate the denoiser threshold on held-out clean tasks at a prespecified false-positive rate, then freeze it. Compare random perturbations, untargeted PGD and adaptive objectives using the source lambda sweep under matched budgets, steps and multiple restarts. Recover the adaptive budget and preprocessing convention rather than guessing them. Measure threshold evasion jointly with clean-reference corruption, independent task-classifier decisions and effective rollout success. Attacks that evade while producing independent harm would falsify the strong harmless-evasion interpretation; lower corruption with no measurable task harm would support its empirical scope. e-algorithme-adaptivee-detectione-verifierse-limitse-repro
8.3 Reading coverage
Visual audit: Viewed all 13 supplied PDF pages, including the title/version block, all five figures, all ten tables, Algorithm 1, limitations, reproducibility statement and references. Cross-checked Figure 2's gradient direction and detector comparison against Eqs. (1)–(2) and Algorithm 1; the algorithm specifies the greater-than flag rule. Figure 3's approximate-equality/inequality markers agree with its caption, while its cosine terminology differs from Eq. (2)'s norm-distance definition. Inspected every final crop at its native rendered resolution, including the narrow adaptive table rendered at 400 DPI. The six assets retain complete panel labels or table headers/rows; no table footnotes are omitted. No appendix appears in this PDF; separate supplements remain unverified. Figure 5 was inspected but is not cropped.
PDF pages inspected for this edition: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13. Appendix coverage: not present.
Original figures and tables remain the work of the source’s authors. Extractions preserve their scientific content; any HTML wrapper layout is disclosed with each figure. The surrounding reading notes are our own.
Text reading scope & known omissions
- Abstract
- 1 Introduction
- 2 Related Work
- 3 Threat Model
- 4 Method: Definitions 1–2, Eqs. (1)–(2), Algorithm 1
- 5 Experiments, including Sections 5.1–5.9
- 6 Mechanism
- 7 The task-level null
- 8 Limitations
- 9 Conclusion
- Reproducibility
- References
Outside the original text pass
- Text extraction does not reconstruct figure images; inspect the retained PDF for figures and equation/table layout.
- Separate supplemental material availability has not been fully verified.
- Separate supplemental material availability has not been fully verified; none was supplied.
- Text extraction does not reconstruct figures; this gap was addressed by viewing every supplied PDF page and all six final crops.
- No appendix is present in the supplied 13-page PDF. Code was not inspected and experiments were not reproduced.
- Identity/version note: title and all three authors match the catalog. The supplied artifact is arXiv:2606.22966v1, stamped 22 Jun 2026, while its title block says June 23, 2026. The catalog submission date is June 22; no different edition was supplied or compared.
The visual audit above records the subsequent illustrated pass.
8.4 Traceable evidence
e-identityPDF p. 1, title block and arXiv margin
The exact catalog title appears with Linghan Chen, Kaiyan Ji, and Minyu Guo, all at Adelaide University. The margin identifies arXiv:2606.22966v1, 22 Jun 2026; the title block is dated June 23, 2026.
Go to primary source ↓e-threatPDF pp. 1–4, Sections 1 and 3, Figure 1 and Table 1
The paper distinguishes a reactive consumer of imagined features from an oracle trusting a predicted future. The attacker knows the encoder and WAM, perturbs the current observation within an infinity-norm budget, and cannot modify weights or access oracle internals or the true future. Only MPC is evaluated end to end.
Go to primary source ↓e-pipelinePDF p. 4, Figure 2, Section 4 and Definition 1
The perturbed observation passes through frozen encoder E and world model W to an imagined latent, then an oracle decision. A backward gradient path updates the observation perturbation; a separate branch supplies the detector. Differentiability is specified for continuous latent targets without discrete VQ or in-place cache writes.
Go to primary source ↓e-objectivesPDF p. 5, Definition 2, Eqs. (1)–(2)
Untargeted loss is negative cosine divergence from the clean imagined latent. Targeted loss is squared norm distance to a chosen imagined target. The displayed gap_closed formula is one minus the adversarial-to-target norm distance divided by the clean-to-target norm distance.
Go to primary source ↓e-algorithmPDF p. 5, Attack and detector paragraph and Algorithm 1, lines 1–9
All parameters are frozen and reset calls use no_grad. Starting at zero, delta follows projected sign-gradient descent with symbolic T and alpha. SPSA is described when a cache path severs gradients. The detector averages future-frame velocity-prediction norms at t in {0.8, 0.9, 0.97} and flags s greater than tau.
Go to primary source ↓e-modelsPDF p. 3, Section 2 first paragraph; p. 6, Table 2
RynnVLA-002 shares a backbone but its action head does not consume imagined tokens. LingBot-VA interleaves video and action tokens with block-causal attention to the generated latent. LaDi-WM predicts DINO/SigLIP latent evolution and refines diffusion-policy actions.
Go to primary source ↓e-protocolPDF p. 5, Section 5 opening; p. 6, Scale at a glance
Experiments use LIBERO-10 and a single-A100 cluster at Adelaide Phoenix. Targeted probes span one to eight tasks, corruption/control five, detection eight, and adaptive attacks four. Reactive results aggregate 143–190 episodes per condition across ten tasks; MPC uses task 4 with 20 episodes per budget. Wilson 95% intervals and Fisher tests are specified.
Go to primary source ↓e-rynnPDF p. 6, Section 5.2
The reported RynnVLA cross-model imagination-attack transfer has 100% task success, matching random and clean, while directly attacking the victim action gives 0%. Same-model gradient coupling and impact-proxy damage do not establish transferred rollout harm.
Go to primary source ↓e-targetedPDF pp. 6–7, Section 5.3, Table 3 rows O1–O3 and Figure 3
Same-observation instruction swapping closes 0.447 of the gap in O1 and 0.46–0.52 in O2. Cross-scene O3 at epsilon 0.2 over eight tasks reports 0.207 ± 0.056 versus random −0.002 ± 0.008, with verifier 0/8 and decode approximately clean. Figure 3 shows frame 12 for task 0 to 3, with the attacked basket-and-cans scene still unlike the kitchen target.
Go to primary source ↓e-corruptionPDF p. 7, Section 5.4 and Table 4; p. 8, Figure 4
Across five LingBot-VA tasks, Table 4 reports corruption 0.135 ± 0.05 versus random 0.002 at epsilon 0.05, labeled approximately 60 times; at epsilon 0.2 it reports 0.266 ± 0.03 versus 0.013. Figure 4 shows visibly altered decoded frame 12 for task 0 at epsilon 0.2.
Go to primary source ↓e-verifiersPDF p. 7, Table 4 and Which verifier, honestly? paragraph
The fidelity verifier thresholds the same clean-reference cosine optimized by the attack; at epsilon 0.1 and tau 0.9 its flip rate is 100% versus 0% random. Independent nearest-prototype verifier V1 has clean accuracy 1.000, 12% untargeted task misclassification versus 0% random, and targeted false acceptance 0% over eight tasks at epsilon 0.2.
Go to primary source ↓e-negativePDF pp. 7–8, Section 5.5 and Table 5
At epsilon 0.1 over five tasks, future-latent divergence is 0.198 ± 0.045, observation-readout divergence 0.0083, reported amplification 24.5 times, and random future divergence 0.004. This compares two latent channels under the same perturbed observation.
Go to primary source ↓e-detectionPDF p. 8, Section 5.6 and Table 6; p. 6, detection sample scope
Denoiser self-consistency yields AUC 0.89, 1.00, 1.00 at epsilon 0.1, 0.2, 0.4, with TPR at 5% FPR of 62%, 100%, 100%. Surface RMS/max-abs is around chance; the detector needs no clean reference. Detection is reported over eight tasks.
Go to primary source ↓e-adaptivePDF p. 9, Section 5.7 and Table 7; p. 6, adaptive sample scope
The adaptive objective is negative corruption plus lambda times detector score. Lambda 0, 0.5, 2, 8 gives corruption 0.279, 0.241, 0.156, 0.064; score 0.562, 0.407, 0.345, 0.337; AUC 1.00, 0.88, 0.31, 0.06. Clean score is 0.369; the sweep covers four tasks.
Go to primary source ↓e-transferPDF p. 9, Section 5.8
A single perturbation is optimized on tasks [0, 3] and frozen for tasks [4, 7]. Held-out corruption is 0.199 versus random 0.039; training corruption is 0.356 versus 0.034. The perturbation magnitude is written as |delta| = 21.25 without an explicit norm or normalization convention there.
Go to primary source ↓e-mpcPDF pp. 9–10, Section 5.9, Table 8 and Figure 5
LaDi-WM sampling MPC evaluates four candidate actions every five steps. On task 4, clean success is 0.55. With 20 episodes per point, random versus adversarial success is 14/20 versus 1/20 at epsilon 0.01, 13/20 versus 1/20 at 0.03, and 8/20 versus 0/20 at 0.06. At 0.01 the reported Wilson intervals are [0.48, 0.85] and [0.01, 0.24], with Fisher p below 10^-4.
Go to primary source ↓e-causalPDF pp. 9–10, Causal isolation paragraph and Table 9
At epsilon 0.06 with 20 episodes per condition, clean/random/full attack/score_only/final_only success is 0.55/0.40/0/0.60/0. Score_only perturbs scoring with clean execution; final_only reverses this. Direct vision and in-act imagination refinement share the corrupted observation and DINO-SigLIP encoder, preventing their separation by this observation-space intervention.
Go to primary source ↓e-mechanismPDF pp. 10–11, Section 6
The authors explain corruption versus steering using a manifold of natural futures, asserting that untargeted damage moves off it while precise targeting must reach another specified point on it. This is the proposed explanation linking visible damage, detection and the adaptive tradeoff.
Go to primary source ↓e-reactivePDF p. 2, Figure 1 caption; p. 6, rollout scope; p. 11, Section 7 and Table 10
LingBot-VA reactive success is 0.978 clean and 0.966 with oracle-attacked imagination at epsilon 0.8; a direct observation attack at epsilon 1.5 gives 0.943. Universal adversarial success spans 0.83–1.0 at epsilon 0.8–1.2 while random falls to zero. These are aggregated reactive rollouts, distinct from the single-task LaDi-WM experiment.
Go to primary source ↓e-limitsPDF p. 11, Section 8
The paper states simulation-only and single-task MPC limits, shared-encoder attribution ambiguity, and bounded cross-scene steering. A safety-gate pilot of three cases is superseded by full-scale N=179: gate AUC 0.771 and effective-success loss 0.18–0.39 across false-positive rates. The semantic CLIP verifier is described as unreliable at the decoded resolution.
Go to primary source ↓e-reproPDF p. 5, Algorithm 1 and experiment setup; p. 12, Reproducibility
The source supplies symbolic PGD steps and step size, frozen/reset handling, and single-A100 hardware description. It says code, attack/defense scripts, decoded frames and per-condition rollout JSONs are available from the authors, without providing a repository link or a complete implementation configuration.
Go to primary source ↓8.5 Primary sources
Attacking the Trusted Imagination: Oracle-Level Integrity Attacks on Imagine-then-Act World Models ↗
PDF · 4,683 extracted words
Source fingerprint
dfaea912e7d9bb39eb7bf5b892c33327d4d4ca9a08d5b083282ef72a1a26e0a3