PAPER REPORTENAll readings ↗

Beyond Task Success: Stage-Wise Reliability of World Model Planning under Sensing Degradation

English reading report: Method, equations, original figures, experiments and reproducibility.

Authors: Geonmyeong Lee; Byoung-Tak Zhang

Affiliations: Seoul National University

Source: 2609.07126 ↗ · Catalog record

Reading: 22 / 558 · 5 original figures & tables · ~19 min ·

1. Paper overview

In one sentence: Paired diagnostics reveal where sensing disturbances attenuate or persist in world-model planning, but internal sensitivity does not by itself establish a change in executed task success. e01e03e04e05e08e09e12

At a glanceWhat to know
Research problem
Source description

Final success hides whether corrupted sensing was suppressed by prediction, changed action preferences, or persisted without changing the binary outcome. The paper asks where disturbances travel through a planner, using controlled clean/degraded comparisons instead of treating one internal score as a complete robustness measure. e01e03e04

Core mechanism
Source description

A four-stage paired evaluation separates representation shift, prediction discrepancy under identical actions, preference changes over identical candidate actions, and independently executed outcomes. e03e04

A key reported resultOGBScene Drawer: attenuation versus executed outcome: Low-light: 1.384 → 0.054 → rank 1, with 52% success (−8 pp). Overexposure: 2.224 → 0.083 → rank 1, with 50% success (−10 pp). Delay-5/10: residuals 0.326/0.705 and ranks 32/89.5, both 66% success (+6 pp).

Stage 1 shift; Stage 2 residual; Stage 3 median rank; Stage 4 success and paired difference.. DINO-WM; 50 paired replan-eligible tasks; goal offset 20.

Clean success is 60%. All primary paired 95% Tango confidence intervals include zero. Internal ordering differs from outcome point estimates. Neither temporal improvement nor equivalence is established. e04e05e07e08

Reading caution
Source description

Controlled synthetic degradations, few task/model settings, and the absence of real sensing validation limit scope. The study does not causally distinguish task-relevant information retention, alternative successful paths, or replanning as explanations for outcome decoupling. e08e14

Core contributions

  • Source description

    A four-stage paired evaluation separates representation shift, prediction discrepancy under identical actions, preference changes over identical candidate actions, and independently executed outcomes. e03e04

  • Reader analysis

    The observed stage orderings and temporal contrasts motivate targeting verification and sensing mitigation at particular pipeline stages. This is a diagnostic contribution, not a new trained planner or demonstrated mitigation algorithm. e07e09e12e14

Figure 1. Four measurement sites distinguish altered sensing from altered behavior. Original paper, p. 2 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Start with panel (a): the upper row changes image appearance, while the lower row changes which observations enter history. In panel (b), follow the arrows from degraded observation/history through the visual encoder and latent representation to action-conditioned future prediction, CEM candidate evaluation, and environment execution. The brackets mark measurement stages rather than independent trained models. The caption specifies that the goal observation remains clean. Sections 3.2–3.3 clarify that prediction and CEM evaluation are coupled during planning; the analysis separates them by holding actions or candidate pools fixed. This distinction is essential when interpreting the apparent sequence of boxes. e02e03e04e05

What it supports. The figure locates the intervention at sensing and the measurements downstream. A large encoder response can be measured independently of whether candidate preference or executed success changes. The scientific contribution is this controlled diagnostic decomposition applied to existing world models, rather than a newly specified neural architecture.

Where the evidence stops. The diagram does not display the replanning feedback loop or detailed neural modules. Its straight arrows should not be read as proof of open-loop execution or independent training of the four stages.

2. Motivation

2.1 The problem and the proposed response

Source description

Final success hides whether corrupted sensing was suppressed by prediction, changed action preferences, or persisted without changing the binary outcome. The paper asks where disturbances travel through a planner, using controlled clean/degraded comparisons instead of treating one internal score as a complete robustness measure. e01e03e04

2.2 What this reading follows

A camera disturbance can alter the representation entering a world model without equally altering its predictions, preferred actions, or eventual task outcome. Lee and Zhang make those distinctions measurable by following the same sensing condition through four stages of an existing planner. Their primary study uses DINO-WM on 50 paired OGBScene Drawer tasks, supplemented by another task and model. The most instructive contrast changes which history frame becomes stale: almost equal initial shifts yield very different prediction and ranking responses. Read the figures as a diagnosis of propagation, with explicit uncertainty at the outcome stage and no demonstrated recovery of subsequent task success. e01e03e04e05e08e09e12

3. Research context

We place the paper in the collection through its world–action interface. The catalog labels and the reading’s assessment are shown separately.

Catalog dimensionRecorded classification
Major categoryNot assigned
ArchitectureNot assigned
Prediction paradigmNot assigned
QuadrantNot assigned

This table preserves the labels recorded at reading time. The current major category is Evaluation metrics. View the current classification.

3.1 Evidence-based assessment

Insufficient evidence to decide

Reader analysis

The recorded taxonomy is entirely unassigned, so there is no substantive catalog label to confirm. Architectural evidence supports action-conditioned forward prediction plus a separate CEM search procedure in the evaluated pipeline. This diagnostic paper does not establish a new One Model architecture, joint future/action generation, or inverse-dynamics control, and it should not receive a quadrant solely from mentioning VLA systems. e01e03e04e05

This is the collection’s architectural analysis, not a new related-work survey. Benchmark comparisons and their protocols appear in Section 6.

4. Problem formulation

4.1 Inputs and outputs

InputsOutputs
  • Clean/degraded visual observation histories and an unchanged clean goal observation.
  • Paired start–goal tasks; shared action sequences for prediction and shared CEM candidates for preference diagnosis.
  • Normalized latent shifts and relative prediction residuals.
  • Clean-winner ranks, same-top counts, task success, and paired success-rate differences with confidence intervals.

4.2 Equations and their role

D5Dcontext\frac{D_5}{D_{\mathrm{context}}}
The paper's unnumbered relative prediction residual: D_context is the initial clean/degraded context discrepancy and D_5 is the reported future-prediction discrepancy under the same actions. Near-zero ratios indicate attenuation; larger ratios indicate persistence. This is discrepancy between model predictions, not prediction error against a ground-truth future. The paper does not fully specify the distance implementation or aggregation. e03e04e15

5. Method in detail

5.1 Separate a changed context from a changed action

Source description

The first tutorial step is to understand what the paired controls remove. If clean and degraded planners chose different actions before prediction was compared, a difference in predicted futures could arise from those actions rather than sensing. Stage 2 instead takes the degraded planner's chosen sequence and supplies it to both contexts. Its residual D_5/D_context measures how much context discrepancy remains relative to the initial discrepancy; it does not measure accuracy against the environment's true future. Stage 3 solves a related confound by scoring one shared pool of 300 iteration-0 CEM candidates under both contexts. Tracking the clean winner's new rank reveals preference change without mixing in independently sampled candidate pools. Stage 4 deliberately changes the protocol again: each planner executes its own selected actions, so task outcomes include the consequences of decision changes. e03e04

Table 1. Corruption type includes information position and age, not just image severity. Original paper, p. 3 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Read each row across from condition to parameter to operation. The visual family includes intensity scaling, directional blur, transient local masking, and droplet-like local blur/refraction. The temporal family is more structured: random frame loss reuses the last valid observation, Prior maps [A,B,C] to [A,A,C], and Current maps it to [A,B,B]. Thus Prior preserves the newest observation C while Current removes it. Delay-5 and Delay-10 shift the entire history to older observations; they are not synonyms for random missing updates. These explicit operators explain why similarly sized representation shifts need not carry the same information into prediction. e02e09e14

What it supports. The table provides concrete perturbation settings, including a 0.40 drop probability and whole-history delays of 5 or 10 environment steps. Its substitution pair creates a targeted comparison of stale information at different history positions while retaining a three-observation context.

Where the evidence stops. These are synthetic operators, not calibrated measurements of real camera failures. Nominal frame-loss probability also differs from actual planner exposure, so it cannot alone describe the corruption consumed during a task.

5.2 Ask which information changed, not just how much

Reader analysis

The Prior/Current pair makes the central temporal argument concrete. Starting from [A,B,C], Prior produces [A,A,C], preserving the newest observation, while Current produces [A,B,B], making that observation stale. Figure 3 reports almost matching representation shifts, yet their prediction residuals and candidate ranks separate. The source associates this contrast with history position and latest-observation freshness. A reader's interpretation is that aggregate latent distance loses information about which part of the context was altered; it cannot alone specify how a predictor will respond. This is not evidence that earlier frames are irrelevant: Prior still changes the top-ranked candidate in 22 of 50 tasks. Random frame loss adds another qualification: corrupted frames entered 46 histories, but the latest observation was dropped in only 21 tasks. Nominal corruption and consumed corruption must therefore remain distinct. e02e09

5.3 Keep diagnostic sensitivity separate from recovery

Reader analysis

Outcome interpretation needs a second boundary beyond experimental pairing. On Scene, the delay conditions strongly change prediction and preference, yet their success-rate point estimates exceed the clean baseline; all paired confidence intervals include zero. The authors discuss task-relevant information, alternative successful paths and replanning as possible explanations, but do not causally separate them. The Cube table then changes the sampling question: it asks what happens to tasks each model already solves, with smaller exposed subsets for delay. Finally, latest-frame restoration changes history at a fixed physical state and reduces distance to the clean decision in a supporting DEV study. It never follows that intervention through a subsequent executed trajectory. A reader should therefore treat restoration as a promising hypothesis about a sensing intervention, while reserving claims of behavioral recovery for an experiment that actually measures paired task outcomes. e06e08e12e13e14

5.4 Training and inference

During training

Source description

DINO-WM supplies a frozen visual representation and learned action-conditioned predictor; LeWM is described as a JEPA-style model with different representations and training. This paper evaluates those models and supplies no new training loss, training schedule, dataset recipe, or compute accounting. e05e15

During inference

Reader analysis

The studied flow encodes observations, predicts action-conditioned futures, evaluates CEM candidates, and executes the chosen action in the environment with replanning. Stages 2 and 3 are separated experimentally although prediction and candidate evaluation are coupled in the planner. The paper does not jointly generate future states and actions or use an inverse-dynamics action decoder. e03e04e05

5.5 Implementation flow

  1. Apply controlled sensing operators

    Table 1 specifies low-light gain 0.20, overexposure factor 2.20 with clipping, directional blur kernel 11, transient 45×45 occlusion, and eight water droplets covering 20%. Temporal conditions drop updates with probability 0.40 and reuse the last valid frame, substitute one history position, or shift the whole history by 5 or 10 environment steps. Only observation/history is corrupted; the goal stays clean. e02

  2. Measure representation sensitivity

    The frozen visual encoder receives clean and degraded inputs. Latent distance is normalized by representation change over one normal clean-trajectory step. For temporal conditions, the additional current-observation distance separates latest-frame change from whole-history shift. e03

  3. Hold actions fixed during prediction

    The degraded planner's selected action sequence is injected into both contexts. Comparing predicted future representations therefore removes action-selection differences from this diagnostic. The relative residual divides future discrepancy by initial context discrepancy. e03e04

  4. Hold the candidate pool fixed during planning

    Both contexts score the same 300 candidate sequences generated at CEM iteration 0, using latent-distance cost to the clean goal. Report the degraded-context rank of the clean winner and same-top: how often that winner remains rank 1. This measures preference, not necessarily final-action or trajectory disagreement. e04

  5. Execute paired outcomes

    Stage 4 independently plans and executes selected actions for each sensing condition on the same start–goal tasks. Success uses the environment's goal criterion within its episode budget. The primary DINO-WM OGBScene Drawer evaluation uses goal offset 20 and 50 fixed replan-eligible tasks. e04e05

6. Experiments & results

This diagnostic study follows synthetic sensing disturbances through an existing world-model planner's representation, action-conditioned prediction, candidate ranking, and executed task outcome. Paired OGBench evaluations show that a disturbance's relative impact can change between stages. Temporal information position matters even when aggregate representation shifts are similar. These diagnostics locate sensitivity; they neither establish a general predictor of task failure nor demonstrate successful sensing mitigation.

Source and visual limitations
Reader analysis

The paper supplies a functional planning pipeline and perturbation/recovery diagnostics rather than a new neural architecture or a training-component ablation. Figure 3 serves as the diagnostic visual; its recovery panel measures decisions at a fixed state, with no visual or experimental evidence of subsequent trajectory or task-success recovery. The edition preserves all three original figures and both tables. e03e04e12e14

6.1 Read the original evidence

Figure 2. Large internal changes and outcome point estimates have different orderings. Original paper, p. 5 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Follow one condition horizontally across four different axes; do not compare horizontal distances between panels as if their units matched. Panel (a) normalizes representation change to a clean trajectory step. Panel (b) reports D_5/D_context under identical actions. Panel (c) gives the median degraded-context rank of the clean winner, with same-top counts in parentheses; preserve the displayed denominators, including 37 for occlusion and 46 for frame loss. Panel (d) plots paired success-rate differences in percentage points against zero, with the clean baseline stated as 60%. The accompanying method identifies its error bars as paired 95% Tango confidence intervals. e03e04e05e07e08e09e10

What it supports. Low-light and overexposure both retain median rank 1 after strong predictive attenuation, yet their reported success rates are 52% and 50%. Delay-5 and Delay-10 produce ranks 32 and 89.5 but both reach 66%. These contrasts illustrate different internal and outcome point-estimate orderings; every primary outcome interval includes zero.

Where the evidence stops. The point estimates establish neither improvement nor equivalence. Also, the Prior/Current outcome ordering disagrees with Figure 3; those condition-specific outcome values are unresolved. The differing Stage 3 denominators should not be silently replaced with 50.

Table 2. Secondary evidence changes appearance sensitivity and exposes selection effects. Original paper, p. 7 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Read within each model's three rows before comparing models. Each was evaluated only on tasks it solved with clean sensing: 17 for DINO-WM and 15 for LeWM. Stage 1 is normalized within the model, so these are not raw distances on one common latent scale. The Stage 3 columns distinguish median rank from same-top count. For Delay-10, the Stage 4 parentheses retain success on the planner-exposed subset: DINO-WM has 9/12 exposed successes alongside 14/17 overall, and LeWM has 2/5 alongside 12/15. The caption explaining these denominators is retained because omitting it would materially change the interpretation. e06e13

What it supports. LeWM's blur row has a small Stage 1 shift, 0.142, yet residual 0.954 and median rank 2; low-light instead gives rank 93 and only 2/15 successes. DINO-WM's corresponding low-light success is 8/17. The secondary study supports heterogeneous propagation and model-dependent appearance sensitivity within the selected tasks.

Where the evidence stops. The models do not share a guaranteed identical task subset. Their success counts are conditional on model-specific clean success, and exposed delay subsets are small. Treat these as secondary diagnostics, not a controlled leaderboard or broad robustness guarantee.

6.2 Results and evaluation conditions

Task & protocolReported resultComparison & interpretation
OGBScene Drawer: attenuation versus executed outcome

DINO-WM; 50 paired replan-eligible tasks; goal offset 20.

Low-light: 1.384 → 0.054 → rank 1, with 52% success (−8 pp). Overexposure: 2.224 → 0.083 → rank 1, with 50% success (−10 pp). Delay-5/10: residuals 0.326/0.705 and ranks 32/89.5, both 66% success (+6 pp).

Stage 1 shift; Stage 2 residual; Stage 3 median rank; Stage 4 success and paired difference.

Clean success is 60%. All primary paired 95% Tango confidence intervals include zero.

Internal ordering differs from outcome point estimates. Neither temporal improvement nor equivalence is established. e04e05e07e08

OGBScene Drawer: history-position diagnostic

Same primary tasks; Substitution-Prior versus Substitution-Current.

Prior: 0.7903, 0.0369, rank 1. Current: 0.7918, 0.5713, rank 19.

Representation shift; relative prediction residual; median clean-winner rank.

Nearly matched aggregate shifts accompany sharply different downstream responses. Prior still displaces the clean winner in 22/50 tasks.

Latest-frame freshness matters, but older history also matters. Condition-specific substitution outcome values remain unresolved because Figures 2 and 3 disagree in ordering. e09e10

OGBScene Drawer: large terminal-error diagnostic

Condition-task cases among clean-successful tasks, pooled separately over five appearance and five temporal conditions.

Large-error cases: 32/150 appearance versus 3/150 temporal. Down-flips: 32/150 versus 4/150; 35/36 down-flips fall among large-error cases.

Terminal-error increase of at least 40 mm; clean-success-to-failure transitions.

Any error increase occurs in 85/150 versus 79/150 cases.

Worsening is concentrated in larger errors, not uniformly more frequent error increases. The 40 mm threshold is post-hoc, not a replacement success rule; these are repeated condition-task cases. e11

Same-state temporal recovery counterfactual

Supporting DEV comparison of stale history, latest observation restored, and clean history; 10 pairs.

7/10 pairs move toward clean; first-action distance decreases 16.63% and full-sequence distance 22.86%.

Movement toward the clean decision; median reduction in decision distance.

Restore only the latest observation while retaining the same physical state.

Partial decision recovery is measured. Subsequent trajectory or task-success recovery is not tested. e12

OGB-Cube: model-specific appearance sensitivity

Secondary evaluation, restricted to each model's clean-solved tasks: DINO-WM 17, LeWM 15.

DINO-WM low-light: 2.46/0.228/2/8-of-17; blur: 4.69/0.104/3/14-of-17. LeWM low-light: 7.13/0.846/93/2-of-15; blur: 0.142/0.954/2/11-of-15.

Stage 1 shift; Stage 2 residual; median rank; success count.

Each model's selected subset is successful under clean sensing.

Appearance sensitivity varies by model. Stage 1 is normalized within each model; model-specific selection prevents a controlled head-to-head success comparison. e06e13

OGB-Cube: Delay-10 exposure

Same secondary clean-solved subsets; distinguish all selected tasks from planner-exposed subsets.

DINO-WM: 0.947, 93.5, 0/12; success 14/17 overall and 9/12 exposed. LeWM: 0.990, 157, 0/5; success 12/15 overall and 2/5 exposed.

Residual; median rank; same-top; success.

Exposure counts differ from total task counts.

Overall success can mask the effect on actually exposed tasks. The small selected subsets limit generalization. e06e13

6.3 Ablations and diagnostic examples

Read component removals and qualitative examples within their stated evaluation conditions.

Figure 3. Similar aggregate history shifts can conceal different downstream sensitivity. Original paper, p. 6 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. In panel (a), check which history slot is replaced and follow the whole-history delay arrows toward older time indices. In panel (b), hollow markers denote Prior or Delay-5 and filled markers denote Current or Delay-10. The Prior/Current Stage 1 values are almost identical, while the Stage 2 residual and Stage 3 rank separate markedly. Panel (c) compares all-stale [S,S,S], latest-restored [S,S,C], and clean [A,B,C] histories, with the displayed ordering S<A<B<C describing freshness. The 16.63% and 22.86% labels mean reductions in distance to the clean first action and entire action sequence, as clarified by the recovery paragraph on page 7. e02e09e10e12

What it supports. Prior versus Current yields residuals 0.0369 versus 0.5713 and median ranks 1 versus 19, despite shifts 0.7903 versus 0.7918. Latest-only restoration moves 7/10 DEV decisions toward clean. Together these observations support investigating information position and freshness, without establishing that only the newest observation matters.

Where the evidence stops. Figure 3 labels Prior +4 pp and Current +2 pp, reversing Figure 2's ordering. Its two delays both show +6 pp, contradicting Section 4.2's phrase 'every stage.' Recovery is a same-state DEV intervention, not evidence of recovered trajectories or success.

7. Analysis & limitations

7.1 What the evidence leaves open

Source description

Controlled synthetic degradations, few task/model settings, and the absence of real sensing validation limit scope. The study does not causally distinguish task-relevant information retention, alternative successful paths, or replanning as explanations for outcome decoupling. e08e14

Reader analysis

Figure 3 labels substitution outcome changes Prior +4 pp and Current +2 pp, whereas Figure 2 places Current to the right of Prior. Their ordering is inconsistent and unresolved. Section 4.2 also says longer delay increases changes at every stage, but Figure 3 shows +6 pp for both delays at Stage 4; only Stages 1–3 support an increase. e10

7.2 Questions for discussion

  1. Which stage metric predicts task-relevant failure after controlling for actual corruption exposure?
  2. Does refreshing the latest observation improve executed success, or merely restore similarity to the clean planner?

8. Reproducibility audit

8.1 Requirements and known gaps

Reader analysis

Reproduction needs the exact model checkpoints, OGBench tasks and seeds, goal-offset configuration, corruption implementations, clean normalization trajectories, matched prediction actions, shared 300-candidate pools, and paired outcome records. Table 1 specifies several corruption strengths, but the paper does not provide task IDs, checkpoint versions, software/hardware versions, complete CEM settings, numerical episode budget, or full distance/aggregation definitions. e02e03e04e05e15

Reader analysis

Proposed checks should first repeat the Prior/Current comparison with recorded frame ages and matched candidate randomness, then execute a controlled latest-frame restoration intervention. Decision-distance recovery and paired task-success recovery must be tested separately. e09e12e14

8.2 Proposed reproduction checks

The following checks are proposals motivated by the paper. They have not been run as part of this reading.

Check 1: Repeat the history-position contrast with exposure controls

Reader-proposed, not performed: at identical Scene states, form clean, Prior and Current histories, reuse matched CEM randomness, and record the exact frame ages consumed. For Stage 2, preserve the paper's matched-action protocol within each clean/degraded pair; for Stage 3, reuse the 300-candidate pool. Report per-task initial discrepancies, residuals, ranks and same-top, alongside a comparison restricted to closely matched initial shifts. Repeat across predefined task and candidate seeds. The freshness interpretation would be weakened if Current's larger residual and rank shift disappear under these controls or are explained by a few unmatched contexts. Preserve separate outcome records to resolve the conflicting substitution labels in Figures 2 and 3. e02e04e05e09e10

Check 2: Test whether latest-frame restoration recovers executed success

Reader-proposed, not performed: fork matched delayed Scene states into continued-stale, latest-only-restored and fully-clean-history conditions. Keep goals, planner settings and remaining episode budgets identical, and hold the future observation policy fixed across branches so the initial history intervention is isolated. First measure distance to the clean first action and full action sequence, then execute each branch and record terminal error and paired success with the paper's paired confidence-interval method. Predefine the task sample and analysis before inspecting outcomes. A repeat of decision-distance reductions without improved executed outcomes would falsify the stronger claim that this restoration alone recovers task performance; success recovery cannot be inferred from the original DEV result. e04e05e12e14

8.3 Reading coverage

Visual audit: All eight supplied PDF pages were rendered and visually inspected, including title/authors/version on p. 1, Figure 1 on p. 2, Table 1 and representation definitions on p. 3, prediction/planning/outcome protocols and evaluation setup on p. 4, Figure 2 and outcome/error analysis on p. 5, Figure 3 and temporal exposure analysis on p. 6, Table 2/recovery/limitations on p. 7, and final references on p. 8. All five final original crops were separately viewed. Figure 1 arrows and temporal operators were checked against Table 1 and Section 3.2; Figure 3 markers and recovery labels were checked against Figures 2–3 and the p. 7 text. The conflicting substitution outcome ordering and overbroad delay prose are disclosed. The short Table 2 caption is retained because it defines normalization and exposed-subset notation. No appendix appears in this PDF; separate supplements and external cited works were not supplied or inspected.

PDF pages inspected for this edition: 1, 2, 3, 4, 5, 6, 7, 8. Appendix coverage: not present.

Original figures and tables remain the work of the source’s authors. Extractions preserve their scientific content; any HTML wrapper layout is disclosed with each figure. The surrounding reading notes are our own.

Text reading scope & known omissions
  • Abstract (p. 1)
  • 1 Introduction (pp. 1–2)
  • 2 Related Work (p. 2)
  • 3 Sensing Degradation and Stage-Wise Evaluation (pp. 3–4), including 3.1–3.3
  • 4 Results (pp. 4–7), including 4.1–4.3
  • 5 Conclusion (p. 7)
  • References (pp. 7–8)

Outside the original text pass

  • Text extraction does not reconstruct figure images; inspect the retained PDF for figures and equation/table layout.
  • Separate supplemental material availability has not been fully verified.
  • Identity: the inspected title, Geonmyeong Lee and Byoung-Tak Zhang, and identifier match the catalog. The title page identifies arXiv:2609.07126v1 [cs.RO], 7 Sep 2026, and Seoul National University. No alternate revision or edition was supplied or compared.
  • All three supplied text chunks and all eight PDF pages were read. The extraction's Figure 1 image-slot placeholders do not represent absent source images: the original PDF images were visually inspected.
  • Separate supplemental material availability has not been fully verified; no separate supplement was supplied. The eight-page PDF contains no appendix.
  • Code and external cited works were not inspected; no experiments were reproduced. No implementation or release availability is inferred from bibliography links.

The visual audit above records the subsequent illustrated pass.

8.4 Traceable evidence

e01PDF p. 1, title/author block, arXiv margin, Abstract and Section 1Inspect

Exact title and authors match the catalog; affiliation is Seoul National University; the artifact is arXiv:2609.07126v1, 7 Sep 2026. The problem is locating sensing disturbance propagation beyond final success.

Go to primary source ↓
e02PDF p. 2, Figure 1 and caption; p. 3, Table 1 and Section 3.1Inspect

Ten corruption operators and parameter values are specified. The goal observation stays clean; temporal operators modify observation history. Water-drop regions remain fixed during an episode.

Go to primary source ↓
e03PDF p. 2, Figure 1(b); p. 3, Section 3.2, opening paragraph and Stages 1–2Inspect

The functional pipeline separates representation, prediction, preference and execution. Stage 1 uses a frozen encoder and clean-step normalization, with a separate current-observation distance for temporal conditions.

Go to primary source ↓
e04PDF p. 4, Section 3.2, Stages 2–4Inspect

Prediction compares identical degraded-selected actions; the residual is D_5/D_context. Planning uses the same 300 iteration-0 candidates and goal-latent cost. Outcomes use independent executions, paired differences and 95% Tango intervals.

Go to primary source ↓
e05PDF p. 4, Section 3.3, Primary Scene evaluationInspect

DINO-WM has frozen visual representation, learned action-conditioned prediction and CEM planning. OGBScene Drawer uses goal offset 20 and the same 50 fixed replan-eligible start–goal tasks.

Go to primary source ↓
e06PDF p. 4, Section 3.3, Secondary Cube and LeWM evaluationInspect

OGB-Cube moves a cube to a 3-D target. DINO-WM and JEPA-style LeWM are tested on low-light, blur and Delay-10, each restricted to its own clean-solved tasks.

Go to primary source ↓
e07PDF pp. 4–5, Section 4.1, internal-stage propagation; p. 5, Figure 2(a–c)Inspect

Low-light and overexposure shifts attenuate to residuals 0.054 and 0.083 and median rank 1. Substitution-Current has shift 0.792, residual 0.571 and rank 19. Overexposure/water-drop same-top counts are 34/50 and 23/50 despite median ranks 1 and 2.

Go to primary source ↓
e08PDF p. 5, Figure 2(d) and Section 4.1, internal-to-outcome decouplingInspect

Clean success is 60%; low-light and blur reach 52%, overexposure 50%, both delays 66%. All primary paired 95% intervals include zero. The experiments do not causally separate proposed explanations for decoupling.

Go to primary source ↓
e09PDF p. 6, Figure 3(a–b) and Section 4.2, temporal profiles and actual exposureInspect

Prior/Current shifts are 0.7903/0.7918, residuals 0.0369/0.5713 and ranks 1/19. Random loss enters 46/50 histories and drops the latest observation in 21 tasks; Prior changes the top candidate in 22/50.

Go to primary source ↓
e10PDF p. 5, Figure 2(d), substitution rows; p. 6, Figure 3(b), Stage 4; Section 4.2, final sentenceInspect

Figure 2 places Current's positive outcome shift farther right than Prior's, but Figure 3 labels Prior 4 and Current 2. Figure 3 shows both delay outcome shifts at 6, despite prose saying delay increases changes at every stage. These source inconsistencies remain unresolved.

Go to primary source ↓
e11PDF p. 5, Section 4.1, appearance–temporal Stage 4 contrast and terminal-error paragraphsInspect

Down-flips are 32/150 versus 4/150; up-flips 15/100 versus 13/100. Large-error cases are 32/150 versus 3/150 and include 35/36 down-flips. Any increase occurs in 85/150 versus 79/150. The 40 mm threshold is a post-hoc diagnostic.

Go to primary source ↓
e12PDF p. 6, Figure 3(c) and caption; p. 7, Recovery paragraphInspect

Latest-only restoration moves 7/10 DEV pairs toward clean; median first-action and sequence distances decrease 16.63% and 22.86%. It is a same-physical-state counterfactual without downstream trajectory or success recovery evidence.

Go to primary source ↓
e13PDF p. 7, Table 2, all six rows and caption; Section 4.3Inspect

Table 2 supplies per-model shifts, residuals, ranks, same-top and success. DINO-WM uses 17 and LeWM 15 clean-solved tasks; Delay-10 parentheses give 9/12 and 2/5 exposed successes. Stage 1 normalization is model-specific.

Go to primary source ↓
e14PDF p. 7, Section 5, final paragraphInspect

The authors limit conclusions to synthetic degradations and a small set of settings, call for broader models/tasks/real sensing, and leave executed recovery from targeted interventions for future work.

Go to primary source ↓
e15PDF pp. 3–4, Sections 3.1–3.3; pp. 6–7, recovery protocol and conclusion; pp. 7–8, end of paperInspect

The supplied methods define diagnostic operations and selected settings but do not include a new training objective, full training recipe, checkpoint/task identifiers, complete planner and metric implementation details, compute/software inventory, or an appendix supplying them.

Go to primary source ↓

8.5 Primary sources

Scroll across the image to inspect details. Press Esc to close.