PAPER REPORTENAll readings ↗

Towards Practical World Model-based Reinforcement Learning for Vision-Language-Action Models

English reading report: Method, equations, original figures, experiments and reproducibility.

Authors: Zhilong Zhang; Haoxiang Ren; Yihao Sun; Yifei Sheng; Haonan Wang; Haoxin Lin; Zhichao Wu; Pierre-Luc Bacon; Yang Yu

Affiliations: National Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China; School of Artificial Intelligence, Nanjing University, Nanjing, China; Mila - Quebec AI Institute; Université de Montréal

Source: ICML 2026 · 2603.20607 ↗ · Catalog record

Reading: 303 / 558 · 6 original figures & tables · ~20 min ·

1. Paper overview

In one sentence: VLA-MBPO improves a separate robot policy with short, multiview imagined rollouts, trading fewer environment interactions for world-model compute and residual prediction bias. e02e03e04e05e08e10e13e14

At a glanceWhat to know
Research problem
Source description

Pixel-input VLAs need imagined observations usable by their visual encoder and rewards accurate enough for policy learning. Independent camera predictions can disagree, while long model rollouts accumulate errors that can reverse sparse success signals. The paper seeks better use of a fixed interaction dataset when collecting additional robot experience is costly. e02e04

Core mechanism
Source description

Adapt a pretrained unified multimodal model to action-conditioned dynamics and reward prediction, retaining its vocabulary; impose head-to-wrist information flow through interleaved decoding. e03e04

A key reported resultLIBERO policy success: Reported Avg 85.9; displayed suite scores: Spatial 87.8, Object 96.6, Goal 92.8, Long 66.8.

Average success rate (%). Four suites, 10 tasks each; 50 offline episodes and 50 evaluation episodes per task; one-trajectory SFT initialization.

Average: SFT 76.8, piRL 82.6, IDQL 77.5, BC(WM) 76.0. Long: SFT 54.6; piRL 61.2. The printed Avg entries imply a 9.1-percentage-point gain over SFT; Long gains 12.2 points. Table 2 has two unresolved inconsistencies: Goal's printed delta is +6.8, although 92.8 minus 85.8 equals 7.0 points; the four displayed VLA-MBPO suite scores average 86.0%, whereas Avg prints 85.9%. These reported values are preserved without silently correcting the source. Equivalent interaction budgets do not establish equal total training compute; no uncertainty is supplied. e09e10

Reading caution
Source description

The authors report high sample-generation cost and continued dependence on downstream action-labeled data. Failure examples include unobserved scene structure, large-motion collapse and implausible towel deformation. e16e21

Core contributions

  • Source description

    Adapt a pretrained unified multimodal model to action-conditioned dynamics and reward prediction, retaining its vocabulary; impose head-to-wrist information flow through interleaved decoding. e03e04

  • Source description

    Combine chunk-level world modeling, short branches from offline observations and Flow-Noise PPO. The authors support this design with value-gap bounds, model-quality tests and executed-policy evaluations. e05e06e08e10e13

Figure 1. One model supplies imagined observations and rewards; a separate VLA learns from short branches. Original paper, p. 2 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Start in panel A with the current camera images and integer action tokens beneath the unified model. The first generated image supplies context for the wrist-view outputs, following the head-to-wrist direction in Eq. (3). The final text question elicits a success answer. In panel B, gray recorded observations provide branch points, while blue squares mark short imagined continuations. Those samples enter an imagined buffer and then Flow-Noise PPO. The return arrow labeled Model Weights summarizes iteration: Algorithm 1 explicitly updates the VLA policy and value head inside the RL loop, after fitting the world model. Panel C provides task examples rather than additional model components. e03e04e05e17

What it supports. The method separates two jobs: generate useful training experience and learn a controller from it. Its unified component combines visual dynamics with reward prediction, while the VLA remains a distinct policy. Short branches let learning revisit intermediate recorded states without asking the world model to simulate every step of an entire task.

Where the evidence stops. The diagram alone does not resolve reward aggregation. Section 3 defines a discounted within-chunk reward, whereas Appendix D supplies a binary success prompt and describes a pretrained VLM. The exact implementation connecting these descriptions remains unspecified.

2. Motivation

2.1 The problem and the proposed response

Source description

Pixel-input VLAs need imagined observations usable by their visual encoder and rewards accurate enough for policy learning. Independent camera predictions can disagree, while long model rollouts accumulate errors that can reverse sparse success signals. The paper seeks better use of a fixed interaction dataset when collecting additional robot experience is costly. e02e04

2.2 What this reading follows

A robot policy cannot benefit from imagined experience unless the images and rewards remain useful for control. VLA-MBPO tackles that requirement in two places. A pretrained unified multimodal model predicts the next head view, conditions wrist predictions on it, and supplies a success signal. A separate VLA then learns through short branches starting from recorded observations, with a critic estimating what lies beyond each branch. The strongest evidence combines LIBERO policy gains, a rollout-length ablation and physical-robot execution. The reading below follows that chain while keeping model fidelity, task success and the limits of the implementation description distinct. e02e03e04e05e08e10e13e14

3. Research context

We place the paper in the collection through its world–action interface. The catalog labels and the reading’s assessment are shown separately.

Catalog dimensionRecorded classification
Major categoryWAMs
ArchitectureDual-system
Prediction paradigmOther mechanisms
QuadrantOutside quadrants

3.1 Evidence-based assessment

Supports the recorded classification

Reader analysis

The recorded Dual-system / Other mechanisms / Outside quadrants classification fits a separate action policy and world model used for policy post-training. Unified dynamics/reward prediction inside UMM-World does not make the entire controller One Model. Multiview conditioning is supported, but the 3D multiview subcategory should not imply an explicit 3D state representation, geometric reconstruction loss or inverse-dynamics action extraction. e03e04e18

This is the collection’s architectural analysis, not a new related-work survey. Benchmark comparisons and their protocols appear in Section 6.

4. Problem formulation

4.1 Inputs and outputs

InputsOutputs
  • Current head and wrist images; task instruction
  • Low-level action chunks discretized into integer text tokens for the world model
  • Recorded trajectories from the initial policy, plus expert data as specified for physical tasks
  • Predicted chunk-end camera observations and a learned reward/success signal
  • An RL-finetuned VLA policy producing low-level action chunks

4.2 Equations and their role

st+khTθ(sth,stw,at:t+k1)st+kwTθ(stw,st+kh)\begin{aligned}s^h_{t+k}&\sim T_\theta(\cdot\mid s^h_t,s^w_t,a_{t:t+k-1})\\s^w_{t+k}&\sim T_\theta(\cdot\mid s^w_t,s^h_{t+k})\end{aligned}
Eq. (3): h and w denote head and wrist observations, t is environment time, k is chunk size, a is the control sequence, and T_theta is the learned transition model. Wrist generation has no direct action argument here; the inspected training mask is consistent with this restriction. e04e18
TtV=j=1kγj1r(st+j,l)+γkVϕ(st+k,l)Vϕ(st,l)\mathcal{T}_t^V=\sum_{j=1}^{k}\gamma^{j-1}r(s_{t+j},l)+\gamma^k V_\phi(s_{t+k},l)-V_\phi(s_t,l)
Eq. (4), residual only: r is the per-step task reward, l the instruction, gamma the discount, and V_phi the learned value. The bootstrap carries consequences beyond the branch. The source does not clearly reconcile this within-chunk reward sum with its binary endpoint prompt. e05e17

5. Method in detail

5.1 Make the imagined observation match the policy's interface

Source description

The VLA expects camera images and language, so a low-dimensional imagined state would not directly replace its observations. Section 3.1 adapts Bagel to generate images after receiving a discretized sequence of controls. The first prediction is the future head view. Eq. (3) then conditions a future wrist view on that predicted head view and its own current wrist view, without a direct action argument. Figure 10's wrist-generation rows mask the action columns and expose the corresponding current wrist and clean predicted head features. This is an information-flow choice, not an explicit reconstruction of 3D geometry. Appendix D distinguishes ViT semantic features from VAE image features; its mask caption uses t for diffusion noise time, whereas Eq. (3) uses t for environment time. Keeping these two meanings separate prevents a misleading temporal reading of the mask. e04e17e18

5.2 Use a critic to connect short imagined segments

Reader analysis

After data collection and world-model fitting, Algorithm 1 repeatedly samples recorded start observations and creates short imagined branches. Eq. (4) combines rewards within an action chunk with a discounted value prediction at its endpoint. Reader interpretation: the critic carries the burden of estimating consequences beyond the branch, so short rollouts exchange some dynamics-model exposure for dependence on value accuracy. Figure 3 shows selected value estimates moving toward full-trajectory returns, which the authors interpret as stitching across chunks. Theorem 4.2 motivates a tradeoff between policy divergence and chunk-model error as branch length changes; it does not establish the accuracy of a learned success detector. The implementation description also leaves open how the binary endpoint prompt supplies the formal discounted reward sum. A reproduction must settle that mapping before treating the displayed residual as executable pseudocode. e03e05e06e11e17

5.3 Test the entire chain from image quality to execution

Reader analysis

Table 1 supports the world-model component: interleaving improves wrist LPIPS and reward F1, although reward accuracy moves slightly in the opposite direction. Table 2 then asks a different question—whether policy optimization improves simulated task completion—and Table 3 shows that a two-chunk branch works better than the tested shorter, longer and full-horizon alternatives. Figure 5 supplies the physical execution test, but merges seen and unseen conditions. Reader interpretation: together these results make a stronger case than image examples alone, while leaving the mediation from improved cross-view fidelity to improved control unisolated. Figure 6 also increases imagined sample count without a compute-matched control. Appendix F reports eight H100 GPUs and hours of model and policy training, so interaction efficiency should be assessed alongside computational cost and the failure examples in Appendix E. e08e10e13e14e15e21e22

5.4 Training and inference

During training

Source description

World-model settings include learning rate 2e-5 and cross-entropy/MSE weights 0.01/1.0. RL uses action chunks of 10, two rollout chunks, three denoising steps, gamma 0.99 and lambda 0.95. Actor/critic learning rates are 5e-6/1e-4. Contrary to the blanket shared-hyperparameter claim, Long uses 1280 samples and 50 updates-to-data versus 512 and 20 elsewhere. e17e19e05

Author claim

Flow-Noise forms the PPO likelihood from the action-denoising chain. Conservative model-bias penalties, rejection sampling and extra Q models are omitted; the authors attribute this simplification to world-model accuracy. e05

During inference

Source description

During imagined training, the policy proposes controls, UMM-World predicts their observation consequences and reward, and the critic supplies a bootstrap beyond the short branch. In physical evaluation the finetuned policy acts from observed images and language; the described algorithm does not introduce inference-time world-model search. e03e04e05e12

Reader analysis

The visual states provide feedback between action chunks. Cross-view agreement can improve the input the policy receives without guaranteeing correct physical dynamics or successful execution. e04e08e21

5.5 Implementation flow

  1. Collect and fit

    Collect trajectories with the initial VLA, then finetune Bagel as UMM-World. Continuous controls become a k-by-d token sequence: k actions, each with d dimensions. The source gives [0,256] as an example discretization range; it does not specify a complete normalization recipe. e03e04e17

  2. Predict a consistent next observation

    Generate the future head view from both current views and the action sequence. Generate each future wrist view using its current view and the predicted head view. The world model skips intermediate video frames; the robot still executes the underlying controls. e04e03

  3. Branch and improve the policy

    Sample start observations from the offline buffer, generate short action-chunk trajectories, and update the VLA and its appended MLP value head using Flow-Noise PPO. Algorithm 1 fits the world model before this loop and lists no world-model update inside it. e03e05

6. Experiments & results

VLA-MBPO finetunes a separate vision-language-action policy using short imagined rollouts from a Bagel-based dynamics/reward model. Its central choices are action-chunk prediction, head-to-wrist conditional generation and branching from recorded observations. Table 2 reports LIBERO average success rising from 76.8% to 85.9%, with arithmetic discrepancies detailed below; physical-task gains accompany substantial compute costs and unresolved implementation details.

6.1 Read the original evidence

Table 1. Interleaving particularly helps wrist-image fidelity, while reward accuracy and F1 tell different stories. Original paper, p. 7 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Read the two image-metric groups separately: lower LPIPS and higher PSNR/SSIM are preferred for both head and wrist views. The reward columns assess task-success prediction, not image quality. Ctrl-World has no reward entries, and Qwen3-VL has no dynamics entries, so dashes are missing capabilities in this comparison rather than zero scores. Compare UMM-World with the row directly below it to isolate removal of interleaved view decoding, then with the last row for removal of pretrained initialization. These results use LIBERO-Object: 50 training trajectories per task and 10 held-out trajectories per task, evaluated over 40-step rollouts across 10 tasks. e07e08

What it supports. Wrist LPIPS improves from 0.454 without IVD to 0.254 with it, while reward F1 rises from 0.799 to 0.861. The full model also improves head/wrist LPIPS over Ctrl-World's 0.150/0.435. However, ACC is slightly higher without IVD, 98.5 versus 98.4; the evidence does not support improvement on every reward metric.

Where the evidence stops. The inference-time column gives 21, 10 and 8 without a unit or a complete timing protocol. The table reports no uncertainty, explicit cross-view geometry metric or downstream policy ablation for IVD; image-quality gains alone do not establish executed-control gains.

Table 2. The largest gain over the initial policy appears on LIBERO-Long. Original paper, p. 8 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. The One-Trajectory SFT line describes initialization, not the full RL data budget. Section 5.2 collects 50 episodes per task and evaluates 50 episodes per task over 10 tasks in each suite. Compare methods vertically within each suite. The blue delta row reports percentage-point gains over SFT, but Goal prints +6.8 although 92.8 minus 85.8 equals 7.0. Likewise, the four displayed VLA-MBPO suite scores average 86.0%, not the printed Avg of 85.9%. The source does not reconcile these inconsistencies; retain its reported values while flagging the arithmetic. BC(WM) imitates successful generated trajectories, IDQL is offline model-free RL, and piRL is online RL under a stated equivalent interaction budget. That budget does not establish compute equivalence. e09e10e19

What it supports. Table 2 reports 85.9% average success for VLA-MBPO, versus 76.8% for SFT and 82.6% for piRL; its printed average is retained with the discrepancy above disclosed. Long scores are 66.8%, 54.6% and 61.2%, respectively, giving a consistent 12.2-point gain over SFT. BC(WM)'s 48.6% on Long shows that imitating successful generated trajectories does not reproduce the reported RL improvement.

Where the evidence stops. The Goal delta and VLA-MBPO average remain unresolved source inconsistencies. No confidence intervals or seed variation accompany the table. Appendix D increases Long's sample and update budgets, qualifying the shared-hyperparameter claim. These are simulated results; physical execution is evaluated separately in Figure 5.

Figure 5. Physical execution improves across five tasks, with seen and unseen conditions pooled. Original paper, p. 9 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Use the legend to compare the yellow SFT bar, light green IDQL bar and dark green VLA-MBPO bar within each task group. Plug In and Fold Towel use the Arx-X5 bimanual platform; Insert Pen, Pick Cup and Wipe Board use Galaxy-R1. Each task is evaluated over 50 trials, comprising 30 seen and 20 unseen conditions. Appendix C specifies 50 plug, 60 towel and 100 per Galaxy-task expert demonstrations, plus 50 autonomous trajectories per task for RL data collection. The plot's Plug In label corresponds to the task called Plug Cable in the main experimental description. e12e13

What it supports. The dark green bar is highest in all five groups, demonstrating improvement in actual robot task completion under the reported protocol. For Wipe Board, the axis suggests approximately 58% success for VLA-MBPO, compared with approximately 36% for IDQL and 28% for SFT. These are readings of unlabeled bars rather than tabulated exact values.

Where the evidence stops. Figure 5 provides aggregate bars without separate seen/unseen scores or error bars. It therefore supports overall improvement but cannot quantify unseen-condition generalization independently. The main text's approximate Arx data counts should be read alongside the more specific appendix counts.

6.2 Results and evaluation conditions

Task & protocolReported resultComparison & interpretation
LIBERO policy success

Four suites, 10 tasks each; 50 offline episodes and 50 evaluation episodes per task; one-trajectory SFT initialization.

Reported Avg 85.9; displayed suite scores: Spatial 87.8, Object 96.6, Goal 92.8, Long 66.8.

Average success rate (%)

Average: SFT 76.8, piRL 82.6, IDQL 77.5, BC(WM) 76.0. Long: SFT 54.6; piRL 61.2.

The printed Avg entries imply a 9.1-percentage-point gain over SFT; Long gains 12.2 points. Table 2 has two unresolved inconsistencies: Goal's printed delta is +6.8, although 92.8 minus 85.8 equals 7.0 points; the four displayed VLA-MBPO suite scores average 86.0%, whereas Avg prints 85.9%. These reported values are preserved without silently correcting the source. Equivalent interaction budgets do not establish equal total training compute; no uncertainty is supplied. e09e10

LIBERO-Object world-model prediction

50 training and 10 held-out trajectories per task over 10 tasks; 40-step test rollouts.

LPIPS 0.094/0.254; F1 0.861.

Head/wrist LPIPS (lower); reward F1 (higher)

Ctrl-World LPIPS 0.150/0.435; Qwen3-VL-8B F1 0.841. Without IVD: wrist LPIPS 0.454, F1 0.799; without pretraining: 0.579, 0.496.

IVD supports image fidelity and F1, but reward ACC slightly increases when IVD is removed (98.4 to 98.5). These metrics measure model predictions, not robot success. e07e08

LIBERO-Long rollout-length ablation

Same suite; one-, two-, four-chunk branches and full-horizon model rollouts.

Two chunks: 66.8.

Success rate (%)

One: 63.9; four: 62.9; full horizon: 52.8.

Two chunks outperform full horizon by 14.0 percentage points. This supports the rollout tradeoff but does not separately measure prediction error or bootstrap bias. e14

Physical Wipe Board execution

Galaxy-R1; 50 trials combining 30 seen and 20 unseen conditions; 100 expert and 50 autonomous training trajectories.

Approximately 58.

Success rate (%), approximate axis readings

Approximately 28 for SFT and 36 for IDQL.

Figure 5 also ranks VLA-MBPO highest on the other four physical tasks. Unlabeled aggregate bars cannot establish separate unseen-condition gains or statistical significance. e12e13

6.3 Ablations and diagnostic examples

Read component removals and qualitative examples within their stated evaluation conditions.

Table 3. Two chunks give the best reported tradeoff between short-branch learning and model error. Original paper, p. 9 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. The numbers 1, 2 and 4 count world-model rollout chunks, not individual robot controls. Table 5 uses 10 actions per chunk, so a two-chunk branch spans 20 controls while requiring two chunk-level transitions. All entries concern LIBERO-Long. Read the branched columns first to see the nonmonotonic dependence on horizon, then compare them with Full Horizon. Section 5.4 interprets very short branches as limiting exploration and trajectory stitching, while longer branches accumulate unreliable predictions. The value bootstrap in Eq. (4) matters here: shortening a branch moves more of the return estimate into the learned critic rather than eliminating the remaining task. e05e06e14e19

What it supports. Two chunks reach 66.8% success, exceeding one chunk's 63.9%, four chunks' 62.9% and full horizon's 52.8%. The 14.0-point advantage over full horizon is consistent with the motivation for branching. The result does not say that the shortest possible rollout is always best, even on the same benchmark suite.

Where the evidence stops. The table does not separately report critic error, reward error, model error or compute-matched rollout budgets. Its pattern supports the proposed tradeoff, but the authors' causal explanation and value-gap analysis are not directly measured by these success rates.

Figure 14. Endpoint prediction can miss large movements and produce implausible deformation. Original paper, p. 27 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. In each example, begin with the image labeled Initial Obs, then compare the orange-bordered reference continuation above with the blue-bordered UMM-World continuation below. The left sequence concerns Wipe Board: focus on the arm configuration and how little it changes between generated future frames. The right sequence concerns Fold Towel: compare towel shape and its relationship to the grippers. Appendix E.2 attributes these errors to difficulty spanning large physical displacement and substantial nonrigid state changes in a single prediction. The figure is a qualitative diagnostic of model behavior, separate from the aggregate executed-policy successes reported earlier. e04e21

What it supports. A short imagined branch can still contain a consequential wrong transition. The authors document motion collapse and physically implausible towel configurations, showing the limits of direct chunk-end prediction. These examples qualify the method's practical gains: reducing the number of generated transitions reduces opportunities for error accumulation but does not make each prediction reliable.

Where the evidence stops. The caption names large prediction horizon, while the accompanying discussion emphasizes movement magnitude and deformation. This is not a controlled horizon sweep at fixed scene and action conditions. The selected examples provide no failure frequency or quantitative link to policy success.

7. Analysis & limitations

7.1 What the evidence leaves open

Source description

The authors report high sample-generation cost and continued dependence on downstream action-labeled data. Failure examples include unobserved scene structure, large-motion collapse and implausible towel deformation. e16e21

Reader analysis

Theorems compare returns under bounded reward and total-variation assumptions. Their model-error quantities refer to different transition models/distributions; smaller coefficients alone do not prove a tighter realized error. Learned reward-classifier error is not separately controlled. e06

Reader analysis

Figure 3 offers selected value-learning traces, not a controlled demonstration that stitching causes policy gains. Quantitative results provide no run-to-run uncertainty; physical results merge seen and unseen conditions. e11e10e13

7.2 Questions for discussion

  1. Does interleaved decoding improve executed-policy success when generation compute is matched?
  2. How sensitive are short-rollout gains to reward calibration and critic bootstrap error?
  3. Do physical gains persist when seen and unseen trials are reported separately?

8. Reproducibility audit

8.1 Requirements and known gaps

Source description

Use the reported LIBERO collection/evaluation protocol and preserve the larger Long optimization budget. For physical work, Appendix C specifies 50 plug, 60 towel and 100 per Galaxy-task expert demonstrations, plus 50 autonomous trajectories each; the main text only approximates the Arx counts. e09e12e19

Source description

The reported compute is eight NVIDIA H100 GPUs, about 7–8 hours for world-model training and 4–6 hours for policy optimization. Reproducing physical timing requires Galaxy-R1 data recorded at 10 Hz despite 100 Hz control, or Arx-X5 control at 15 Hz. e20e22

Open question

Resolve action normalization, trainable/frozen module boundaries, software versions and exact training schedules before implementation. Eq. (4) leaves its GAE residual index fixed across the sum; the reward-sum/binary-prompt mapping and the appendix wording about a pretrained VLM also need clarification against the unified-model description. e04e05e17e18e19

8.2 Proposed reproduction checks

The following checks are proposals motivated by the paper. They have not been run as part of this reading.

Check 1: Does interleaved decoding improve control at matched budgets?

Reader-proposed, not run: train pretrained UMM-World with and without IVD on the same LIBERO-Object trajectories, preserving initialization, action encoding and loss settings. Evaluate the same 100 held-out trajectories over 40 steps, measuring per-view LPIPS, reward F1 and an explicitly annotated rate of contradictory gripper/object configurations across views. Then finetune identical VLA initializations with two-chunk branches and equal generated-sample and optimizer-update budgets; evaluate 50 episodes per task. Report generation wall time separately because IVD changes decoding cost. If image metrics improve but policy success does not across repeated seeds, the claimed downstream benefit of view consistency is weakened even though Table 1's perceptual gain is reproduced. e04e07e08e09e17e19

Check 2: Separate rollout error from reward and bootstrap error

Reader-proposed, not run: on LIBERO-Long, record simulator states alongside the offline images to permit paired simulator and world-model continuations. Hold the fitted model, initial VLA, chunk size and optimization settings fixed; compare one-, two- and four-chunk branches while matching total generated transitions and policy updates. First document how success labels become the discounted chunk reward. For identical action sequences, measure endpoint image error, disagreement between predicted and simulator reward sums, and critic error against longer simulator returns. Repeat policy evaluation across seeds. If longer branches lose success without increasing measured model/reward error, or if the ranking disappears after matching budgets, compounding error alone is insufficient to explain Table 3. e04e05e06e09e14e17e19

8.3 Reading coverage

Visual audit: Visually inspected the title/author/version page; all Figures 1–14 and Tables 1–5; Algorithm 1; claim-relevant equations; and the appendix pages supporting hardware, data, prompts, configurations, failure analysis and compute. All six final original crops were separately viewed. Table 2's Goal delta and VLA-MBPO average were checked against the displayed scores; both unresolved arithmetic inconsistencies are disclosed while preserving the printed values. Figure 1's generation direction was checked against Eq. (3) and Figure 10's mask; the mask caption's noise timestep was distinguished from environment time. Figure 14's horizon wording was qualified against the accompanying motion/deformation discussion. All seven text chunks were read, including references and complete Appendices A–F; this visual pass is not independent certification of every proof line. No supplemental material, code or experiments were inspected beyond the supplied PDF.

PDF pages inspected for this edition: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 16, 19, 20, 21, 22, 23, 24, 25, 26, 27. Appendix coverage: reviewed.

Original figures and tables remain the work of the source’s authors. Extractions preserve their scientific content; any HTML wrapper layout is disclosed with each figure. The surrounding reading notes are our own.

Text reading scope & known omissions
  • Abstract and Sections 1–2: motivation, language-conditioned MDP and PPO
  • Sections 3.1–3.3: action-conditioned world model, interleaved decoding, branched rollout and policy optimization
  • Section 4: both value-gap theorems
  • Sections 5.1–5.4: model evaluation, simulation, physical tasks and ablations
  • Sections 6–7, limitations, impact statement and references
  • Appendices A–B: lemmas and proofs
  • Appendix C: robot hardware, task assets and evaluation protocols
  • Appendix D: prompts, attention mask and hyperparameters
  • Appendices E–F: visual examples, failure analysis and computational resources

Outside the original text pass

  • Identity/version note: reviewed arXiv:2603.20607v1 [cs.RO], dated 21 March 2026, with a March 24, 2026 preprint footer. Title and the catalog author string match; the supplied BibTeX reverses Haoxin Lin and Zhichao Wu. The catalog venue says ICML 2026 while its BibTeX says ICLR Workshop; neither venue attribution is established by this preprint. No other revision was supplied or compared.
  • Text extraction does not reconstruct figure images; inspect the retained PDF for figures and equation/table layout.
  • The extraction limitation above was addressed by inspecting all 14 figures, all five tables and claim-relevant equation pages in the supplied PDF. All seven supplied text chunks were individually read, including references and appendices. Reading the proofs does not constitute independent mathematical certification.
  • Separate supplemental material availability has not been fully verified.
  • No code, checkpoints or external resources were inspected; no experiments were reproduced.

The visual audit above records the subsequent illustrated pass.

8.4 Traceable evidence

e01PDF p. 1, title, author/affiliation block, arXiv margin and preprint footerInspect

The title and nine authors match the catalog author string. The artifact identifies arXiv:2603.20607v1, 21 March 2026, and a preprint date of March 24, 2026.

Go to primary source ↓
e02PDF pp. 1–2, Abstract and Section 1Inspect

The paper targets expensive VLA reinforcement learning through pixel-space world modeling, cross-view consistency and shorter imagined rollouts under sparse rewards.

Go to primary source ↓
e03PDF p. 2, Figure 1 A–C; p. 5, Algorithm 1, lines 1–9Inspect

A unified dynamics/reward model supplies imagined trajectories to a separate VLA. Data collection and world-model finetuning precede a loop that updates the policy and value function.

Go to primary source ↓
e04PDF p. 3, Figure 2 and Section 3.1; p. 4, action representation, Eq. (3) and Section 3.2Inspect

Bagel receives discretized action chunks and predicts chunk-end observations without generating intermediate frames. The head prediction conditions on current views and actions; the wrist prediction conditions on its current view and predicted head view. Rollouts branch from offline observations.

Go to primary source ↓
e05PDF p. 3, Eqs. (1)–(2); pp. 4–5, Section 3.3, Eqs. (4)–(5) and Algorithm 1Inspect

Flow-Noise PPO trains the action-chunk policy with an appended MLP value head. Eq. (4) uses discounted within-chunk rewards and a terminal value bootstrap; its printed GAE sum leaves the residual subscript fixed at t.

Go to primary source ↓
e06PDF pp. 5–6, Theorems 4.1–4.2, Eqs. (6)–(7); p. 16, Lemma A.3; pp. 19–20, Theorem B.2 and Eq. (35)Inspect

The branched bound contains policy-divergence terms decaying with rollout length and n times the chunk-model divergence. The proof uses a bounded common reward and total-variation assumptions; it does not separately bound learned reward-classifier error.

Go to primary source ↓
e07PDF pp. 6–7, Section 5.1, Benchmark and BaselinesInspect

World models train on 50 trajectories per LIBERO-Object task and are evaluated over 40 steps on 10 held-out trajectories per task, across 10 tasks. Ablations remove interleaved decoding or pretrained initialization.

Go to primary source ↓
e08PDF p. 7, Table 1, all rows and metric headersInspect

UMM-World head/wrist LPIPS are 0.094/0.254; Ctrl-World has 0.150/0.435. Reward F1 is 0.861 versus Qwen3-VL-8B 0.841. Removing IVD gives wrist LPIPS 0.454, F1 0.799 and ACC 98.5 versus 98.4; removing pretraining gives wrist LPIPS 0.579 and F1 0.496. Inference-time entries 21, 10 and 8 have no stated unit.

Go to primary source ↓
e09PDF p. 7, Section 5.2, Benchmark and BaselinesInspect

Each LIBERO suite contains 10 tasks, with 50 collected episodes per task from a one-shot-SFT policy and 50 evaluation episodes per task. Comparators are SFT, BC on successful model trajectories, online piRL under an equivalent interaction budget, and offline IDQL.

Go to primary source ↓
e10PDF p. 8, Table 2, suite columns, Avg column and delta rowInspect

VLA-MBPO reports 87.8/96.6/92.8/66.8 across Spatial/Object/Goal/Long and 85.9 average. SFT average is 76.8; piRL 82.6; IDQL 77.5; BC(WM) 76.0. Long values are SFT 54.6, piRL 61.2, IDQL 52.2, BC(WM) 48.6. The printed delta row is +9.6/+8.0/+6.8/+12.2/+9.1. Reader arithmetic identifies two source inconsistencies: Goal 92.8 minus SFT 85.8 is 7.0, not the printed 6.8; the displayed VLA-MBPO suite scores average 86.0, not the printed 85.9. The supplied text does not reconcile these differences.

Go to primary source ↓
e11PDF p. 7, Generalizable Credit Assignment; p. 8, Figure 3Inspect

Selected failure/success trajectories compare ground-truth return in blue with predicted value in orange across three displayed training iterations. The authors attribute improvement to cross-chunk value stitching.

Go to primary source ↓
e12PDF p. 9, Section 5.3, Experiment Setup; pp. 20–21, Appendix C.1, task and data protocolsInspect

Five physical tasks use Arx-X5 and Galaxy-R1; each evaluation has 30 seen and 20 unseen trials. Appendix C specifies 50 plug, 60 towel, and 100 per Galaxy task expert trajectories, plus 50 autonomous trajectories per task, with expert data also included for RL.

Go to primary source ↓
e13PDF p. 9, Figure 5, legend, success-rate axis and five task groupsInspect

VLA-MBPO has the tallest success-rate bar on Plug In, Fold Towel, Insert Pen, Pick Cup and Wipe Board. Wipe Board bars are approximately 58% for VLA-MBPO, 36% for IDQL and 28% for SFT. Bars have no numerical labels, uncertainty or separate seen/unseen groups.

Go to primary source ↓
e14PDF p. 9, Section 5.4 and Table 3; p. 10, Rollouts SchemeInspect

LIBERO-Long success is 63.9, 66.8 and 62.9 for one-, two- and four-chunk branching, and 52.8 for full-horizon rollouts. The discussion attributes the tradeoff to stitching/exploration versus compounding prediction error.

Go to primary source ↓
e15PDF p. 10, Figure 6 and Imagined Sample SizeInspect

The plotted success rate increases across generated sample sizes 128, 256, 512, 1280 and 2560. The plot reports no uncertainty or compute-matched comparison.

Go to primary source ↓
e16PDF p. 11, Section 7, LimitationsInspect

The authors identify substantial sample-generation cost and the need for downstream world-model finetuning because the UMM was not pretrained on action-labeled robotic data.

Go to primary source ↓
e17PDF p. 23, Appendix D, Prompt Engineering and Table 4Inspect

The world-model prompt supplies camera observations and integer action arrays, optionally a future head view. The reward prompt requests a binary Yes/No success decision and calls the detector a pretrained VLM. Table 4 lists learning rate 2e-5, cross-entropy weight 0.01 and MSE weight 1.0.

Go to primary source ↓
e18PDF p. 23, Appendix D, Causal Attention Mask Implementation; p. 24, Figure 10 and captionInspect

The mask distinguishes ViT semantic features and VAE reconstruction features; caption t denotes noise timestep, with t=0 noiseless. Wrist-generation rows expose the matching current wrist and clean predicted head features; action columns are masked for those rows.

Go to primary source ↓
e19PDF p. 25, Table 5, all hyperparameter rows; p. 23, HyperparametersInspect

The common settings include chunk size 10, two rollout chunks, three denoising steps, gamma 0.99, GAE lambda 0.95, actor/critic learning rates 5e-6/1e-4. Long changes sample size from 512 to 1280 and Update to data from 20 to 50.

Go to primary source ↓
e20PDF p. 20, Appendix C.1, Galaxy-R1 System and Arx-X5 SystemInspect

Galaxy-R1 has 21 DoF, head and two wrist cameras, 100 Hz control and 10 Hz recorded data. Arx-X5 has 14 DoF, three D435i cameras and 15 Hz control.

Go to primary source ↓
e21PDF pp. 26–27, Appendix E.2, items 1–2 and Figures 13–14Inspect

The authors report partial-observability failures, missing unseen scene structure, motion collapse for large displacements and implausible towel configurations. Figure 14 shows qualitative examples without failure frequencies.

Go to primary source ↓
e22PDF p. 27, Appendix F, Computational ResourcesInspect

All experiments used eight NVIDIA H100 GPUs; reported world-model training takes about 7–8 hours and policy optimization about 4–6 hours under the standard interaction budget.

Go to primary source ↓

8.5 Primary sources

Scroll across the image to inspect details. Press Esc to close.