PAPER REPORTENAll readings ↗

HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone

English reading report: Method, equations, original figures, experiments and reproducibility.

Authors: Simple AI; core contributors: Yuteng Wei, Jinming Ma, Jiawei Wang, Weitao Zhou, Yushen Zuo, Ke Rui, Minglei Li; contributors: Jinhao Zhang, Zhikang Pan, Xiang Wang, Haoran Jia, Huan Du, Zicheng Zeng, Jun Ma, Guiyu Qin, Di Zhang, Xiaofei Li

Source: 2607.25895 ↗ · Project page ↗ · Catalog record

Reading: 111 / 558 · 6 original figures & tables · ~19 min ·

1. Paper overview

In one sentence: A wearable capture-and-curation pipeline supplies deployable task-specific supervision, with near-teleoperation aggregate success obtained from substantially more robot-free demonstrations. e-capturee-curatione-designe-regimese-vla-resultse-wam-resultse-limits

At a glanceWhat to know
Research problem
Source description

Can handheld demonstrations supply the entire target-task post-training signal for a deployable robot policy? HiFi-UMI addresses trajectory drift, relative-hand geometry, timing, and visual coverage together. The claim concerns task-specific supervision; existing foundation checkpoints retain their previous training histories. e-identitye-capturee-limits

Core mechanism
Source description

A head-mounted stereo-inertial rig, hand marker cubes, full-palm grippers, and shared hardware trigger provide portable synchronized capture, evaluated through existing policy families. e-capturee-design

A key reported resultFour-task real-robot post-training: VLA backbones: StarVLA UMI 51.3% (82/160); OpenPI UMI 77.5% (124/160).

Binary full-task success. Wiping, folding, insertion, and sorting; 40 trials/task/policy, 160 aggregate. UMI: 3,200 demonstrations/task outside the evaluation scene; teleoperation: about 300 in-scene.

Teleoperation: 53.8% (86/160) and 74.4% (119/160), respectively; UMI differences −2.5 and +3.1 percentage points. Descriptive aggregate parity under unequal data volumes. Success requires complete, stable completion before timeout without intervention. e-benchmarke-regimese-designe-vla-results

Reading caution
Reader analysis

Four tabletop tasks, one matched gripper/camera interface, and scene shift leave other embodiments and shifts untested. Zero-robot post-training does not exclude robot data in public checkpoint histories or real-robot validation. e-benchmarke-traininge-limits

Core contributions

  • Source description

    A head-mounted stereo-inertial rig, hand marker cubes, full-palm grippers, and shared hardware trigger provide portable synchronized capture, evaluated through existing policy families. e-capturee-design

  • Source description

    The full processed corpus has 20,000+ hours, 4,320,000+ episodes, and 480+ scenes. The reported HiFi-UMI-2K release contains 2,000 hours, 482,100+ episodes, and 110+ scenes, with six-view video, trajectories, gripper states, language, and subtask boundaries. The paper specifies CC BY 4.0 and face masking. e-data

  • Author claim

    UMI-only post-training approximately matches teleoperation across two VLAs and LingBot-VA; separate StarVLA pre-training improves downstream performance. These are within-backbone comparisons, not an architecture ranking. e-designe-vla-resultse-wam-resultse-pretrain-online

Figure 3. Shared head-relative localization connects portable sensing to bimanual trajectories. Original paper, p. 7 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Start with the six thumbnails at upper left: the stereo head pair occupies one column, while the two views from each hand occupy the other columns. The central grippers carry marker cubes, and the lower-right photograph identifies the head cameras, encoders, and synchronization hardware. Section 3.1 explains the geometry: estimate a global head trajectory, locate both hands relative to that head, and compose the poses. The lower-left legend separates head, left-hand, and right-hand paths. The handwriting callouts illustrate local reconstruction detail. These panels describe capture; policy training and deployment use only the four wrist views, as specified in Section 5. e-capturee-interfacee-reconstructione-fidelity

What it supports. The useful design choice is a common observation frame for both hands, alongside wider wrist coverage and synchronized sensing. It supports native relative-hand geometry without instrumenting each collection site with base stations. The collage illustrates the system, while the fidelity table supplies the numerical evaluation.

Where the evidence stops. The graphic says “Drift-free long-horizon” trajectories. Section 3.3.2 instead reports centimeter-level long-horizon drift with local consistency and no global loop closure. Read the figure with that qualification; its handwriting example is not a global accuracy test.

2. Motivation

2.1 The problem and the proposed response

Source description

Can handheld demonstrations supply the entire target-task post-training signal for a deployable robot policy? HiFi-UMI addresses trajectory drift, relative-hand geometry, timing, and visual coverage together. The claim concerns task-specific supervision; existing foundation checkpoints retain their previous training histories. e-identitye-capturee-limits

2.2 What this reading follows

HiFi-UMI asks whether improving handheld demonstration fidelity can remove teleoperation from target-task post-training. Its intervention begins before learning: both hands are localized through a shared head-camera frame, sensors are hardware synchronized, and reconstructed trajectories are checked by simulation replay. The resulting episodes train reactive action policies and a future-predicting world-action model. Read the evidence in that order: capture quality, executed robot success, then diagnostics that isolate only part of the learning system. The comparisons preserve each backbone’s native interface and report small aggregate data-source gaps, while leaving equal-sample efficiency and the causal contribution of individual fidelity improvements unresolved. e-capturee-curatione-designe-regimese-vla-resultse-wam-resultse-limits

3. Research context

We place the paper in the collection through its world–action interface. The catalog labels and the reading’s assessment are shown separately.

Catalog dimensionRecorded classification
Major categoryDatasets
ArchitectureNot applicable
Prediction paradigmNot applicable
QuadrantNot applicable

3.1 Evidence-based assessment

Supports the recorded classification

Reader analysis

Dataset and collection-interface classification is supported by the hardware, curated release, and data-source interventions. HiFi-UMI is not itself a new One Model WAM. LingBot-VA is an evaluated consumer with future-latent-conditioned inverse dynamics; that mechanism should not be assigned to the dataset. e-datae-designe-wam-method

This is the collection’s architectural analysis, not a new related-work survey. Benchmark comparisons and their protocols appear in Section 6.

4. Problem formulation

4.1 Inputs and outputs

InputsOutputs
  • Capture: six camera streams, head/hand IMUs, gripper encoders, and operator-marked boundaries.
  • Policy: four wrist views, language instruction, and backbone-native state/history; head views are excluded.
  • Curated, synchronized and annotated bimanual demonstration episodes.
  • Action chunks specifying relative end-effector poses and absolute gripper openings; LingBot-VA also predicts future video latents.

4.2 Equations and their role

ΔTt0,hj,(m)=(Tt0j)1Tt0+δh(m)j\Delta T_{t_0,h}^{j,(m)}=\left(T_{t_0}^{j}\right)^{-1}T_{t_0+\delta_h^{(m)}}^{j}
Equation (2): T is an end-effector pose, j the arm, m the backbone, t₀ the observed chunk anchor, and δ the native future offset at index h. All rows share that anchor. e-interface
pθ(at:t+H1,zt+1:t+Kht,)=pθ(at:t+H1zt+1:t+K,ht,)pθ(zt+1:t+Kht,)p_\theta(a_{t:t+H-1},z_{t+1:t+K}\mid h_t,\ell)=p_\theta(a_{t:t+H-1}\mid z_{t+1:t+K},h_t,\ell)\,p_\theta(z_{t+1:t+K}\mid h_t,\ell)
Equation (5): a denotes actions, z encoded visual latents, hₜ the video–action history, and ℓ the instruction; H and K delimit predicted action and latent sequences. Future prediction conditions inverse-action decoding at inference. e-wam-method

5. Method in detail

5.1 First make a trustworthy label, then train a policy

Reader analysis

The capture system has two different visual jobs. Wrist views record the interaction a deployed policy will see, while the head stereo pair supplies a shared reference for reconstructing both hands. Offline SLAM uses observations from later in the demonstration, and fiducial markers attach the hand trajectories to that head estimate. This extra information is legitimate during label production; the head pair is excluded from policy inputs. Reconstructed trajectories then face simulation replay, followed by annotation and targeted human verification. Reader interpretation: the pipeline reduces several routes by which an image could be paired with an inaccurate or infeasible action label. However, approximately 96% basic validity describes passage through reconstruction and replay gates, not a 96% probability of correct annotation or successful learned robot behavior. e-capturee-reconstructione-curatione-interface

5.2 Keep a chunk’s coordinate frame fixed until execution

Reader analysis

Equation (2) gives every predicted pose the same measured anchor: the end-effector pose when the observation was taken. A later row is therefore a target relative to that observation, not an increment to add to the previous row. The robot restores each target independently before sending timestamped commands to its controller. This distinction matters even though the text sometimes calls actions relative increments. Translation, Rotation6D, and absolute gripper targets form 20 active bimanual channels, but each backbone packs and samples them differently. StarVLA consumes a prefix before replanning; LingBot-VA executes its native blocks and refreshes causal context with real observations and action history. Reader inference: changing the anchor convention could invalidate a data-source comparison even if both models appear to output compatible tensors. e-interfacee-deploymente-benchmark

5.3 Separate source replacement, initialization, and oracle decoding

Reader analysis

The paper asks three questions with different controls. Replacing teleoperation with UMI fixes the model and deployment stack within each backbone, while accepting the collection pipelines’ different demonstration counts and scenes. Changing StarVLA initialization fixes task-specific UMI data and attributes the subsequent gain to additional pre-training. Finally, substituting true future latents in LingBot-VA bypasses one predictive component to inspect pose decoding. These interventions support different conclusions. The first demonstrates a practical route to deployment; the second gives controlled evidence for a useful initialization; the third localizes what the decoder can do with privileged information. Reader interpretation: none alone proves that microsecond synchronization caused parity, nor that video generation causes most WAM failures. Those causal claims require the additional controlled checks below. e-regimese-designe-pretrain-onlinee-oracle-protocole-limits

5.4 Training and inference

During training

Source description

StarVLA trains end to end with AdamW. Its 4,000-hour pre-training uses batch 2,048 for 180,000 steps plus 5,000 decay steps; post-training uses 50,000 steps and batch 512. Its baseline shares Qwen weights but starts with a random action head. OpenPI initializes from pi05_base and uses continuous flow matching without auxiliary text or subtask losses. e-traininge-vla-methode-design

Source description

LingBot-VA initializes from lingbot-va-base, updates the Transformer, freezes the causal VAE and text encoder, and trains 3,500 steps per task/source with batch 32 and equally weighted video/action flow losses. Checkpoints are selected by fixed validation and frozen before final evaluation. e-traininge-wam-method

During inference

Source description

VLAs synchronize observations using calibrated sensor latency and restore each target independently from the query-time pose. StarVLA uses eight Euler steps, predicts 20 actions, and executes ten before replanning; OpenPI uses ten Euler steps and its native horizon. e-vla-methode-deployment

Source description

LingBot-VA executes 12 actions after reset and 24 thereafter. Real observations sampled every three executed actions and action history update a rolling KV cache. Inference is synchronous at chunk boundaries. Pose targets feed 125 Hz inverse kinematics and 1 kHz joint commands; generated video itself does not execute actions. e-deploymente-benchmark

5.5 Implementation flow

  1. Recover both hands in a shared frame

    Offline stereo-inertial SLAM estimates the head trajectory; marker cubes locate both hands in the same head frame. Two non-parallel fisheyes per hand cover about 200 degrees horizontally and vertically. A GPIO trigger aligns cameras, IMUs, and encoders; online warnings flag poor capture. e-capture

  2. Reconstruct, replay, annotate, and curate

    Offline optimization uses future observations and sliding-window local consistency, omitting global loop closure in changing scenes. Abnormal estimates are recomputed. Reconstruction and simulation replay each pass about 98%, giving approximately 96% cumulative basic validity. AI drafts multi-view annotations; humans preferentially verify flagged or uncertain samples. Export balances content and retains quality metadata. e-reconstructione-curation

  3. Preserve the action coordinate contract

    Every future pose is relative to the same observed chunk-start pose. Each arm has three translation, six Rotation6D, and one absolute-gripper channel: 20 active channels, packed into native 20/32/30-dimensional tensors. Validity filtering precedes train/validation splitting; normalization remains condition-specific. e-interfacee-filter

  4. Use distinct policy consumers

    StarVLA combines Qwen3-VL-4B-Instruct with a flow-matching action DiT, adding self-attention after odd-numbered blocks. OpenPI uses PaliGemma and a Gemma action expert. LingBot-VA predicts future latents before inverse-dynamics action decoding in a block-causal video–action sequence. e-vla-methode-wam-method

6. Experiments & results

HiFi-UMI combines wearable robot-free capture with offline reconstruction, simulation replay, and curation to produce action-grounded demonstrations. Across three policy backbones, task-specific UMI post-training reaches approximately the aggregate success of teleoperation under the tested collection regimes. This supports a practical data pipeline, with substantially more UMI demonstrations and matched grippers and wrist cameras.

Source and visual limitations
Source description

The paper contains an initialization/data-count ablation and an oracle pose-decoding diagnostic, but no controlled degradation of pose accuracy, inter-gripper geometry, synchronization, or field of view. The included diagnostic visuals therefore do not isolate the individual hardware fidelity contributions. e-insertione-oracle-protocole-limits

6.1 Read the original evidence

Table 2. Capture fidelity is measured in several distinct units and at different stages. Original paper, p. 12 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Read the description column before comparing values. Pose accuracy is a local end-effector measurement, while synchronization is a cross-sensor timing offset; neither is a policy prediction error. The frame-drop row states its six-camera, 25-fps setting. Reconstruction counts trajectories passing SLAM, and gripper error measures opening angle. The accompanying Section 3.4 specifies that the 3 mm mean translation error is measured against base-station tracking over roughly 2 m of accumulated head motion. That reference system is used for evaluation only. Simulation replay is a second gate described elsewhere: its approximately 98% pass rate combines with reconstruction to yield about 96% basic validity. e-fidelitye-curatione-reconstruction

What it supports. The table reports a locally precise, tightly aligned source of action labels: 3 mm pose error and timing offsets below 40 microseconds. These measurements make the hardware contribution concrete. The 98% reconstruction figure indicates pipeline throughput rather than the probability that a learned robot completes a task.

Where the evidence stops. The table supplies no uncertainty intervals or detailed sample count for its local accuracy test. Simulation replay checks feasibility; it does not certify contact-rich task success. Long-horizon drift remains larger than the local 3 mm figure.

Figure 9. UMI and teleoperation yield small aggregate gaps in both VLA comparisons. Original paper, p. 21 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Compare solid and hatched bars of the same color first: blue denotes StarVLA-QwenPI and red denotes OpenPI. Solid bars use UMI post-training; hatching denotes teleoperation. Each task contributes 40 real-robot trials, and the shaded aggregate pools 160 per policy. The protocol requires full task completion before timeout, a stable final state, and no safety intervention. UMI supplies 3,200 demonstrations per task from other scenes; teleoperation supplies about 300 from the evaluation environment. Consequently, the paired bars address whether the two practical data pipelines can produce similar deployment outcomes. The larger color-to-color differences do not isolate the value of a data source. e-vla-resultse-benchmarke-regimese-designe-limits

What it supports. StarVLA reaches 51.3% with UMI versus 53.8% with teleoperation. OpenPI reaches 77.5% versus 74.4%, reversing the direction of the aggregate gap. Its UMI insertion result is 85.0%. Across these comparisons, removing target-task teleoperation does not produce a consistent aggregate penalty.

Where the evidence stops. One trial changes a task rate by 2.5 percentage points. These bars have no success-rate confidence intervals and do not establish statistical equivalence. Sample volume and scene diversity differ between data sources, so equal-sample efficiency remains untested.

Figure 11. A future-predicting policy also transfers from UMI task supervision to robot execution. Original paper, p. 23 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Use the solid–hatched pairing within each row to compare LingBot-VA’s post-training sources. Both variants start from the same base checkpoint and retain the same video–action objective and execution schedule. Their model predicts future visual latents before decoding actions, but these plotted rates measure complete real-robot tasks. The shaded aggregate summarizes four equally sized sets of 40 trials. Folding ties, wiping and sorting modestly favor UMI, and insertion favors teleoperation. Read this chart independently from the VLA absolute rates: the paper retains the WAM’s own temporal and training protocol rather than imposing a common model cadence across families. e-wam-resultse-designe-wam-methode-deploymente-benchmarke-oracle-protocol

What it supports. UMI achieves 91 successes out of 160, versus 92 for teleoperation: 56.9% and 57.5%. The different signs of the task gaps support the paper’s descriptive aggregate-parity reading. Insertion remains a weakness for the UMI variant, reaching 42.5% against 50.0% for teleoperation.

Where the evidence stops. This is the full generated-future control pipeline, unlike the following oracle diagnostic. The paper’s descriptions of smoother motion and repeated grasp attempts are qualitative; they establish neither faster completion nor a measured cause of the insertion gap.

6.2 Results and evaluation conditions

Task & protocolReported resultComparison & interpretation
Four-task real-robot post-training: VLA backbones

Wiping, folding, insertion, and sorting; 40 trials/task/policy, 160 aggregate. UMI: 3,200 demonstrations/task outside the evaluation scene; teleoperation: about 300 in-scene.

StarVLA UMI 51.3% (82/160); OpenPI UMI 77.5% (124/160).

Binary full-task success

Teleoperation: 53.8% (86/160) and 74.4% (119/160), respectively; UMI differences −2.5 and +3.1 percentage points.

Descriptive aggregate parity under unequal data volumes. Success requires complete, stable completion before timeout without intervention. e-benchmarke-regimese-designe-vla-results

Four-task real-robot post-training: LingBot-VA

Same four tasks/data regimes; matched WAM execution within backbone; 40 trials/task/source.

UMI 56.9% (91/160).

Binary full-task success

Teleoperation 57.5% (92/160); −0.6 percentage points.

The aggregate gap is one rollout. Insertion favors teleoperation, 50.0% versus 42.5%; reported motion smoothness is qualitative, not timed efficiency. e-benchmarke-regimese-wam-results

Remote Insertion: UMI data scaling

OpenPI with 400/800/1,600/3,200/6,400 UMI episodes; 40 randomized robot trials/setting.

37.5% / 65.0% / 70.0% / 85.0% / 82.5%.

Task success

Separate StarVLA ablation: UMI-pretrained initialization at 800 episodes reaches 57.5%, versus 52.5% for Qwen-initialized random action head at 3,200.

The one-trial decline at 6,400 does not establish harm. Quarter-data efficiency is specific to this task and initialization. e-insertione-pretrain-online

WAM pose decoding with oracle future video

Four equally weighted tasks; held-out real-robot episodes, ground-truth future latents, full first 12-action chunk; grippers excluded.

UMI→Real: 24.33 [22.48, 26.32] mm; 0.65 [0.58, 0.72] degrees.

XYZ trajectory RMSE (mm); adjacent-frame rotation error (degrees), with 95% episode-bootstrap intervals

Real→Real: 21.64 [21.19, 22.09] mm; 0.46 [0.43, 0.49] degrees. UMI→UMI: 21.13 mm; 0.88 degrees.

Measures pose decoding with oracle conditioning; it does not identify video prediction as the dominant deployment bottleneck. Rotation measures local increments, not absolute orientation. e-oracle-protocole-oracle-results

StarVLA offline pre-training transfer

Fixed 4,000-hour mixture; fixed held-out chunks and ten separately collected unseen UMI tasks; checkpoint comparison.

Held-out error falls 61%; mean unseen-task error falls 41%. Pre-decay held-out fit: α=0.268, R²=0.993.

Relative action-MSE reduction; exposure-fit exponent

First versus later checkpoints; all unseen tasks improve, with slower cloth-folding transfer.

Exposure scaling on a fixed corpus, not dataset-size scaling or robot success on those ten tasks. The coverage explanation is an author interpretation. e-offline

Four-task StarVLA transfer from UMI pre-training

Same 3,200 UMI episodes/task and post-training recipe; only initialization changes; 40 robot trials/task.

69.4% with UMI pre-training.

Aggregate full-task success

51.3% from Qwen-VL plus random action head; reported gain +18.1 percentage points.

Supports a useful initialization on this backbone. OpenPI in the figure is contextual, not the controlled comparator. e-designe-pretrain-online

6.3 Ablations and diagnostic examples

Read component removals and qualitative examples within their stated evaluation conditions.

Figure 10. Additional task data and a pretrained initialization are two separate interventions. Original paper, p. 22 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. The horizontal axes count task-specific UMI episodes. Panel (a) holds OpenPI’s initialization and recipe fixed while increasing that count. Panel (b) uses StarVLA with additional UMI pre-training; its dashed line is a different initialization trained with 3,200 task episodes, not a second scaling curve. The annotation points to 800 episodes because the pretrained model first exceeds the fixed reference there. Every point is evaluated with 40 robot trials. Read the two panels as complementary experiments rather than one head-to-head comparison: the left examines task-data quantity, and the right asks whether a learned visual-motor initialization reduces the task data needed for this insertion setting. e-insertione-pretrain-onlinee-limits

What it supports. OpenPI improves from 37.5% at 400 episodes to 85.0% at 3,200, then reaches 82.5% at 6,400. StarVLA with UMI pre-training reaches 57.5% at 800 episodes, above the 52.5% reference using four times as many task episodes. The pretrained curve itself is non-monotonic.

Where the evidence stops. The one-trial OpenPI decline and small StarVLA differences should not be read as precise saturation thresholds. The ablation concerns one task and does not degrade hardware fidelity or match UMI against teleoperation at equal counts.

Figure 12. Oracle future video tests the pose decoder while bypassing future-generation errors. Original paper, p. 25 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Read each arrow as training source followed by evaluation domain. Real→Real and UMI→Real share held-out robot observations; UMI→UMI uses episode-disjoint UMI observations. Every learned row receives cached ground-truth future latents. Panel (a) scores XYZ RMSE across the complete first 12-action chunk; panel (b) scores adjacent-frame rotation increments for steps 2 through 12. Both axes are logarithmic and lower is better. Points average source episodes within tasks and weight the four tasks equally. Whiskers are 95% intervals from 20,000 stratified episode-bootstrap resamples. The lower rows sample independent random action channels before applying the same normalization inversion and physical reconstruction. e-oracle-protocole-oracle-resultse-interface

What it supports. UMI→Real reaches 24.33 mm XYZ RMSE and 0.65 degrees local rotation error, compared with 21.64 mm and 0.46 degrees for Real→Real. These learned decoders are far below the random references. The result supports useful cross-domain pose decoding when the future visual transition is supplied.

Where the evidence stops. Oracle future video is unavailable during deployment, and this experiment does not establish that future generation dominates closed-loop failures. Gripper channels are excluded because acquisition semantics differ. Sub-degree adjacent-frame error is not sub-degree absolute orientation accuracy.

7. Analysis & limitations

7.1 What the evidence leaves open

Reader analysis

Four tabletop tasks, one matched gripper/camera interface, and scene shift leave other embodiments and shifts untested. Zero-robot post-training does not exclude robot data in public checkpoint histories or real-robot validation. e-benchmarke-traininge-limits

Reader analysis

No controlled degradation isolates synchronization, pose, relative-hand geometry, or field of view. Unequal demonstration counts prevent an equal-sample efficiency claim. Forty trials give 2.5-percentage-point task resolution; success-rate confidence intervals and a formal equivalence test are not supplied. e-limitse-vla-resultse-wam-results

Reader analysis

Figure 3 labels trajectories drift-free, but Section 3.3.2 reports centimeter-level long-horizon drift. The 3 mm measurement is local and uses external tracking only as evaluation ground truth. e-capturee-reconstructione-fidelity

7.2 Questions for discussion

  1. Would post-training parity persist with matched sample counts and scene diversity?
  2. Which fidelity axis changes closed-loop success when data quantity and content remain fixed?
  3. How much action error appears when LingBot-VA replaces oracle future latents with generated latents?

8. Reproducibility audit

8.1 Requirements and known gaps

Reader analysis

Reproduction needs calibrated transforms, identical grippers/wrist cameras, per-condition normalization, split manifests, and native loaders/executors. Training schedules are given, but accelerator counts/models, training duration, numerical filtering thresholds, exact split sizes, and numerical OpenPI deployment horizon/replanning interval are unspecified. e-interfacee-filtere-traininge-deploymente-benchmark

Reader analysis

The 2,000-hour release is smaller than the 4,000-hour experimental training mixture. Exact subset membership, annotation-model configuration, and replay-controller settings need verification. The edition proposes controlled timing degradation and oracle-versus-generated WAM comparisons. e-datae-curatione-traininge-oracle-protocole-limits

8.2 Proposed reproduction checks

The following checks are proposals motivated by the paper. They have not been run as part of this reading.

Check 1: Degrade synchronization while fixing data and compute

Reader-proposed experiment, not performed: choose one fixed UMI post-training subset for Remote Insertion and Shirt Folding. Train matched policies with original alignment and with deliberately imposed image-to-action offsets, for example ±10 and ±40 ms. Keep episode identities, camera coverage, transforms, normalization, initialization, optimization steps, and deployment timing identical; restrict all variants to the shared temporal support so no condition gains extra samples. Evaluate on the same randomized initial-condition bank with multiple training seeds and report success uncertainty and contact/regrasp failures. A reproducible deterioration with increasing offset would support a timing contribution. Similar outcomes would weaken a timing-specific explanation within the tested range; they would not refute the combined capture system. e-capturee-fidelitye-regimese-designe-limits

Check 2: Measure the cost of replacing true future latents

Reader-proposed offline diagnostic, not performed: for identical held-out real-robot source episodes and a frozen UMI-trained LingBot-VA checkpoint, compare cached true future latents, model-generated future latents, and shuffled true latents as a negative control. Hold conditioning history, action noise seeds, full 12-step horizon, normalization, shared-anchor reconstruction, and pose scoring fixed. Exclude grippers as in the paper, aggregate by source episode, and bootstrap paired differences with equal task weights. A clear generated-versus-true error increase would quantify the extra difficulty associated with generated conditioning; a small difference would weaken that bottleneck hypothesis. Shuffling tests whether the decoder meaningfully uses the future. None of these offline outcomes alone demonstrates closed-loop success. e-wam-methode-deploymente-oracle-protocole-oracle-results

8.3 Reading coverage

Visual audit: Visually inspected the title/version page, contributor credits, all 15 figures and all four tables, together with the supporting method, training, execution, evaluation, diagnostic, scaling, and limitation pages. All six final original crops were viewed individually. The Figure 3 drift-free label is qualified against Section 3.3.2. No appendix occurs in the supplied PDF; pp. 4–6 and 31–33 were read as text. Separate supplements, linked code, dataset files, and experiments remain outside this pass.

PDF pages inspected for this edition: 1, 2, 3, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30. Appendix coverage: not present.

Original figures and tables remain the work of the source’s authors. Extractions preserve their scientific content; any HTML wrapper layout is disclosed with each figure. The surrounding reading notes are our own.

Text reading scope & known omissions
  • 1 Introduction (pp. 1–4)
  • 2 Related Work, including 2.1–2.3 (pp. 4–6)
  • 3 Data Collection and Processing Pipeline, including 3.1–3.4 (pp. 6–12)
  • 4 Dataset and Release (pp. 12–13)
  • 5 Baselines and Training Setup, including 5.1–5.7 (pp. 12–17)
  • 6 Experiments, including 6.1–6.3 and all diagnostic protocols (pp. 17–28)
  • 7 Discussion and limitations (pp. 28–29)
  • 8 Conclusion (p. 29)
  • Author Contributions (p. 30)
  • References (pp. 30–33)

Outside the original text pass

  • Text extraction does not reconstruct figure images; inspect the retained PDF for figures and equation/table layout.
  • Separate supplemental material availability has not been fully verified.
  • Identity: the exact title and identifier match. The title page identifies arXiv:2607.25895v1 [cs.RO], 28 July 2026, with collective author Simple AI. Page 30 verifies the individual contributors represented in the catalog; the catalog punctuation and reversed collective name are parsing artifacts. No revision or edition difference is established by the supplied material.
  • The catalog affiliation string “Website: Dataset:” is not an affiliation. The PDF supplies no separate institutional affiliation list; none is inferred.
  • Text extraction does not reconstruct figure images; inspect the retained PDF for figures and equation/table layout. This acquisition omission was addressed by inspecting all 15 figures and all four tables in the PDF.
  • Separate supplemental material availability has not been fully verified. No separate supplement was supplied.
  • All 11 supplied text chunks were read individually, covering every PDF page. No appendix is present in the 33-page PDF. Visual inspection covered pp. 1–3 and 7–30; introductory/related-work pp. 4–6 and reference-only pp. 31–33 were read as text.
  • Code, dataset files, external links, and base-checkpoint histories were not inspected. No training, robot experiments, or reproduction checks were run.

The visual audit above records the subsequent illustrated pass.

8.4 Traceable evidence

e-identityPDF p. 1, title block and arXiv margin; p. 30, Author ContributionsInspect

Title matches the supplied identity; the margin specifies arXiv:2607.25895v1 [cs.RO], 28 July 2026. The title-page author is Simple AI. Page 30 lists seven core contributors and ten additional contributors, matching the individual names in the catalog. No separate institutional affiliation list appears.

Go to primary source ↓
e-capturePDF pp. 7–8, Figure 3 and Sections 3.1.1–3.1.4Inspect

Head stereo-inertial SLAM and fiducial marker cubes compose global hand trajectories from shared head-relative poses. Each hand has two non-parallel fisheyes with about 200-degree horizontal/vertical coverage; the rig includes IMUs, encoders, a full-palm glove, and GPIO synchronization. Online quality warnings and temporal slicing assist capture. Figure 3 labels head/left/right trajectories and uses the phrase drift-free.

Go to primary source ↓
e-reconstructionPDF p. 10, Section 3.3.2 and Figure 5Inspect

Offline optimization accesses future observations, forgoes global loop closure in changing scenes, and uses local consistency in a dynamic sliding window. The text reports centimeter-level long-horizon drift and millimeter-level local accuracy. Abnormal trajectories are recomputed; reconstruction passes 98%.

Go to primary source ↓
e-curationPDF pp. 9–11, Sections 3.3.1–3.3.6 and Figure 4Inspect

Six textual stages cover capture/upload, reconstruction/cleaning, simulation retargeting, AI annotation, human verification, and analysis/export. Replay passes 98% of reconstructed trajectories, yielding about 96% cumulative basic validity. Annotations include task/subtask language, boundaries, objects, anomalies, and uncertainty. Human verification is targeted and sampling-based. Versioned export supports balancing and traceable quality metadata; detailed annotation-model/controller configurations are not provided.

Go to primary source ↓
e-fidelityPDF pp. 11–12, Section 3.4 and Table 2, all rowsInspect

Table 2 reports 3 mm local end-effector error, timing offset below 40 microseconds, fewer than one dropped frame per 270,000 frames for six cameras at 25 fps, 98% reconstruction, and opening-angle error below 0.1 degrees. The 3 mm mean translation error is measured against base-station tracking over about 2 m accumulated head trajectory; external tracking is used only for this evaluation.

Go to primary source ↓
e-dataPDF pp. 12–13, Section 4, Figure 6 and Table 3Inspect

The full processed corpus exceeds 20,000 hours and 4.32 million episodes across 480+ scenes. The curated release lists 2,000 hours, 482,100+ episodes, 110+ scenes, and six camera views per episode. Exports contain synchronized video, bimanual trajectories, gripper states, language, and boundaries. Section 4 specifies CC BY 4.0 and masking all human faces before distribution.

Go to primary source ↓
e-interfacePDF pp. 13–14, Section 5, Equations (1)–(2); p. 15, Section 5.3Inspect

Policies use four wrist views; head stereo is reserved for reconstruction. Each future pose is relative to the current observation anchor, with translation in that end-effector frame, Rotation6D, and absolute gripper targets. Twenty active dimensions are used directly by StarVLA, padded into OpenPI's 32 channels, and mapped into LingBot's 30 channels. LingBot uses the first two rotation-matrix rows and masks unused action channels.

Go to primary source ↓
e-vla-methodPDF p. 14, Sections 5.1–5.2, Equations (3)–(4)Inspect

StarVLA uses Qwen3-VL-4B-Instruct and a conditional flow-matching DiT, cross-attending to all 36 Qwen layers and adding action self-attention after odd blocks. It uses eight Euler steps, a 20-action horizon, ten executed actions, and 224×224 images. OpenPI uses PaliGemma/Gemma joint attention, serialized state, a continuous flow-matching-only objective, and ten Euler steps; text/subtask auxiliary objectives are omitted.

Go to primary source ↓
e-wam-methodPDF p. 15, Section 5.3 and Equation (5)Inspect

LingBot-VA models future latents conditioned on video–action history and language, then actions conditioned on those latents. A block-causal mask orders the video/action chunks. Video and action flow losses are equally weighted. Deployment uses rolling KV context, video/action guidance 5/1, denoising steps 8/16, and attention window 24.

Go to primary source ↓
e-filterPDF pp. 15–16, Section 5.5Inspect

Backbone-specific loaders construct anchored future targets after frame conversion. Episodes with tracking failure, missing frames, discontinuities, pose outliers, physical-threshold action spikes, incomplete execution, or inconsistent gripper states are removed. Filtering precedes splitting; idle boundaries are trimmed and normalization is separated by channel type and arm. Numerical rejection thresholds and exact split sizes are not supplied.

Go to primary source ↓
e-trainingPDF p. 16, Section 5.6Inspect

StarVLA is optimized end to end with AdamW. Pre-training uses 4,000 hours, global batch 2,048, 180k steps and a 5k-step decay export; post-training uses batch 512 and 50k steps. OpenPI starts from pi05_base. LingBot starts from lingbot-va-base, freezes VAE/text encoders, updates the Transformer for 3,500 steps with batch 32 and equal video/action losses. Validation includes real-robot metrics and checkpoints are frozen before final blind evaluation. Accelerator counts/models and training duration are not reported.

Go to primary source ↓
e-deploymentPDF p. 17, Sections 5.7.1–5.7.2 and Equation (6)Inspect

VLAs align observations to current time minus maximum calibrated sensor latency, restore each row from a shared query pose, and buffer timestamped targets. OpenPI retains native horizon/replanning without numerical values here. LingBot executes 12 actions after reset, then 24, with shared chunk anchors; observations every three actions and preceding action history update its rolling cache. Inference is synchronous and cache resets with episode/prompt changes.

Go to primary source ↓
e-benchmarkPDF pp. 17–18, Sections 6.1.1–6.1.2; p. 20, Section 6.1.4Inspect

The platform has two seven-joint Tianji Robotics Marvin M6 arms and capture-matched grippers/four wrist cameras; policies receive no head views. Targets and inverse kinematics run at 125 Hz with 1 kHz EtherCAT joint commands. Evaluation uses two identically configured platforms, separate policy/scene operators, randomized order and initial placements, and 40 trials/task/policy. Success requires the whole objective before timeout, stability for two seconds, and no intervention; partial completion fails. Non-policy hardware faults are repeated. Tasks are Stain Wiping, Shirt Folding, Remote Insertion, and Produce Sorting.

Go to primary source ↓
e-regimesPDF pp. 18–19, Section 6.1.3Inspect

Each task uses 3,200 UMI trajectories (roughly 10–20 hours) from multiple operators/sites versus about 300 teleoperation trajectories (roughly 3–7 hours) collected in the evaluation environment. No UMI trajectory is collected in that scene. Counts reflect practical throughput rather than a sample-matched design.

Go to primary source ↓
e-designPDF pp. 19–20, Table 4 and Section 6.1.4Inspect

C1/C2, C3/C4, and C5/C6 compare post-training sources with initialization, recipe, and deployment fixed within each backbone. C1/C7 compare StarVLA initialization at matched UMI post-training data. Only C7 receives additional HiFi-UMI pre-training; OpenPI and LingBot retain public base checkpoints. Cross-backbone temporal/training settings are not forced to match.

Go to primary source ↓
e-vla-resultsPDF pp. 20–21, Section 6.2.1 and Figure 9, task and aggregate barsInspect

StarVLA UMI/teleoperation aggregate successes are 82/160 (51.3%) and 86/160 (53.8%). OpenPI reaches 124/160 (77.5%) and 119/160 (74.4%). Remote Insertion is 52.5/50.0 for StarVLA and 85.0/77.5 for OpenPI. OpenPI ties at 65.0 for wiping, has 77.5/80.0 for folding and 82.5/75.0 for sorting. Each task has 40 trials; differences are interpreted descriptively.

Go to primary source ↓
e-insertionPDF p. 22, Figure 10 and Section 6.2.1; p. 23, opening paragraphInspect

OpenPI insertion success at 400/800/1,600/3,200/6,400 UMI episodes is 37.5/65/70/85/82.5 percent. StarVLA with UMI pre-training reaches 47.5/57.5/80/75/77.5 percent, versus a dashed 52.5% Qwen-initialized reference using 3,200 episodes. Each setting uses 40 robot rollouts. The authors describe a plateau near 3,200 at this resolution.

Go to primary source ↓
e-wam-resultsPDF pp. 23–24, Figure 11 and Section 6.2.2, aggregate/task-level comparisonInspect

LingBot UMI/teleoperation aggregate successes are 91/160 (56.9%) and 92/160 (57.5%). Wiping is 62.5/60.0, folding 65/65, insertion 42.5/50, and sorting 57.5/55 percent. Larger, more continuous UMI motions and grasp-retry difficulties are qualitative observations; completion time was not a quantitative comparison.

Go to primary source ↓
e-oracle-protocolPDF pp. 24–25, Section 6.2.2, ground-truth-video analysis, Equations (7)–(8), Figure 12 captionInspect

Cached true future latents replace generated latents for an oracle diagnostic. Full H=12 chunks are de-normalized and reconstructed identically. XYZ RMSE averages squared error over six coordinates and H steps; rotation averages adjacent-frame SO(3) error for steps 2–12. Learned records and three inference seeds are averaged per episode, then per task with equal task weights; 20,000 stratified source-episode bootstrap resamples produce 95% intervals. Random references average 256 samples per record. Grippers are excluded because empty-grasp acquisition semantics differ.

Go to primary source ↓
e-oracle-resultsPDF p. 25, Figure 12, learned-policy and true-random rows; accompanying comparisonInspect

UMI→Real XYZ RMSE is 24.33 [22.48,26.32] mm and local rotation error 0.65 [0.58,0.72] degrees; Real→Real is 21.64 [21.19,22.09] mm and 0.46 [0.43,0.49] degrees; UMI→UMI is 21.13 [20.02,22.31] mm and 0.88 [0.84,0.92] degrees. True-random actions yield 117.57 mm/126.47 degrees on real and 123.80 mm/126.49 degrees on UMI. Random channels are sampled independently from Uniform[-1,1] in task-specific UMI normalization space before the same physical scoring.

Go to primary source ↓
e-offlinePDF pp. 26–27, Section 6.3, Figures 13–14 and Equation (9)Inspect

On a fixed 4,000-hour corpus, held-out error falls 61% during the reported pass; the pre-decay exposure fit has α=0.268 and R²=0.993. Ten separately collected tasks are absent from pre-training; mean OOD error falls 41%, with improvement on each task. Figure 14 compares 5k and final exports and groups utensil/tableware, granular transfer, and cloth folding. The authors associate slower cloth transfer with sparse textile data. This is exposure scaling, not dataset-size scaling.

Go to primary source ↓
e-pretrain-onlinePDF p. 27, Section 6.3, benefits for post-training; p. 28, Figure 15Inspect

The controlled StarVLA comparison fixes 3,200 UMI episodes/task, splits, normalization, recipe, and deployment, changing initialization. Success rises from 51.3% to 69.4% aggregate, reported as +18.1 percentage points. Task values change from 57.5/52.5/52.5/42.5 to 77.5/75/75/50 percent for wiping/folding/insertion/sorting. Each task uses 40 rollouts. OpenPI is only a cross-architecture reference.

Go to primary source ↓
e-limitsPDF pp. 28–29, Section 7, empirical implications and limitationsInspect

Zero-robot refers to target-task post-training and does not characterize the prior training of public base checkpoints. Evidence spans four bimanual tabletop tasks and three backbones; additional pre-training is tested only on StarVLA. Forty-trial resolution limits per-task conclusions. Hardware fidelity factors are not individually degraded; data-source comparisons are not sample matched, and broader tasks, embodiments, corpus scaling, and transfer to teleoperation post-training remain open.

Go to primary source ↓

8.5 Primary sources

Scroll across the image to inspect details. Press Esc to close.