Data Pyramid for Embodied Manipulation: A Survey
1. Paper overview
In one sentence: Scalable embodied learning needs complementary data sources whose supervision is aligned to physical execution; the survey maps those tradeoffs without establishing a universally best mixture. e02-pyramide04-umie05-humane06-simulatione07-generale17-mixture-gap
| At a glance | What to know |
|---|---|
| Research problem | Source description Internet-scale perception and language supervision do not automatically identify how a particular robot changes physical state. Robot experience couples observations, actions, and consequences, but collection requires hardware, resets, labor, and supervision. The survey asks how sources with different cost, fidelity, and embodiment dependence should be collected, compared, and combined. Its distinction between scalability and robot alignment makes the missing supervision in each source explicit. e02-pyramide22-survey |
| Core mechanism | Source description A category-level pyramid uses scalability and robot alignment as its principal axes, supplemented by quality, diversity, reusability, and physical fidelity. The ordering is an overall synthesis, not a monotonic ranking on every property. e02-pyramid |
| A key reported result | Real-robot dataset scale and coverage inventory: RoboMIND 2.0: 310K trajectories, 1000+ hours, 739 tasks, 6 embodiment types. Trajectories; hours; tasks; embodiment types. Survey Table 1; published dataset statistics, not a common train/test benchmark or policy evaluation. RoboMIND: 107K trajectories, 305.5 hours, 479 tasks, 4 embodiment types. The inventory documents expanded scale and coverage. It supplies no controlled downstream success comparison, uncertainty estimate, or attribution of improvement to dataset size. e19-robot-statistics |
| Reading caution | Reader analysis The pyramid is qualitative. Figure 4 uses heuristic interaction keyframes as a spatial-diversity proxy, without demonstrating that larger scatter predicts transfer. Different tasks, workspace bounds, and extraction choices prevent interpreting these panels as a controlled ranking. e02-pyramide18-diagnostic |
Core contributions
- Source description
A category-level pyramid uses scalability and robot alignment as its principal axes, supplemented by quality, diversity, reusability, and physical fidelity. The ordering is an overall synthesis, not a monotonic ranking on every property. e02-pyramid
- Source description
The survey connects collection interfaces, available labels, and action representations to three model roles: embodied brains, VLAs, and WAMs. Its model-recipe inventory describes source membership without establishing causal gains from each source. e08-recipese10-braine11-vlae12-wam
- Source description
The forward agenda covers tactile sensing, failure and recovery, scalable collection, cross-embodiment alignment, dexterous transfer from human video, and principled mixtures. These are research directions rather than validated prescriptions. e14-tactilee15-failuree16-transfere17-mixture-gap
Figure 1, left pyramid subpanel. A map of data sources and their relationship to robot execution. Original paper, p. 1 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. Start at the apex and ask what each downward step removes from collection. Real-robot data records actions on the platform that executes them. UMI removes the robot from collection while retaining gripper motion. Ego data retains real human interaction but loses native robot controls. Simulation restores controllable actions and privileged labels while approximating physics. General data broadens semantic and perceptual coverage. The introduction identifies scalability and robot alignment as the main organizing axes, supplemented by quality, diversity, reusability, and physical fidelity. This excerpt is the conceptual pyramid within the larger organizational figure; the photographs illustrate source types rather than a common experiment. e02-pyramide22-survey
What it supports. The central tradeoff concerns the kind of supervision available, not simply the amount of data. Simulation can have executable action labels while human video has real physical interaction; neither dominates every dimension. The authors explicitly describe the ordering as an overall synthesis of several properties.
Where the evidence stops. Layer widths are schematic, not measured dataset volumes, costs, or capability scores. This is a taxonomy figure, not the architecture of a newly proposed policy, and it specifies no optimal training proportions.
2. Motivation
2.1 The problem and the proposed response
Internet-scale perception and language supervision do not automatically identify how a particular robot changes physical state. Robot experience couples observations, actions, and consequences, but collection requires hardware, resets, labor, and supervision. The survey asks how sources with different cost, fidelity, and embodiment dependence should be collected, compared, and combined. Its distinction between scalability and robot alignment makes the missing supervision in each source explicit. e02-pyramide22-survey
2.2 What this reading follows
A robot can watch a person pour a drink without learning the actuator commands or contact forces needed to do it itself. This survey organizes that gap as a data pyramid: real-robot trajectories supply direct control experience, UMI demonstrations preserve end-effector motion, human recordings broaden interaction coverage, simulation supplies controllable synthetic experience, and general data contributes perception and reasoning priors. Read the pyramid as a map of complementary supervision. The tables describe what has been collected and combined; they do not rank policies under a common test. The practical question is which missing signal each additional source supplies, and what alignment is still required before execution. e02-pyramide04-umie05-humane06-simulatione07-generale17-mixture-gap
3. Research context
We place the paper in the collection through its world–action interface. The catalog labels and the reading’s assessment are shown separately.
| Catalog dimension | Recorded classification |
|---|---|
| Major category | Foundational work |
| Architecture | Not applicable |
| Prediction paradigm | Not applicable |
| Quadrant | Not applicable |
This table preserves the labels recorded at reading time. The current major category is Related resources. View the current classification.
3.1 Evidence-based assessment
Supports the recorded classification
The recorded foundational-work classification and survey/data-collection subcategories fit the paper's actual contribution. Architecture, prediction paradigm, and quadrant are correctly Not applicable: this survey compares multiple mechanisms without proposing its own world-action architecture. Discussion of joint video/action models does not make the survey itself One Model, and its attention to touch does not establish a tactile policy contribution. e22-surveye12-wame14-tactile
This is the collection’s architectural analysis, not a new related-work survey. Benchmark comparisons and their protocols appear in Section 6.
4. Problem formulation
4.1 Inputs and outputs
| Inputs | Outputs |
|---|---|
|
|
5. Method in detail
5.1 Follow a UMI demonstration all the way to an executed action
The UMI layer becomes clearer when collection and execution are traced separately. During collection, a person moves a portable gripper while cameras, inertial tracking, and gripper sensing recover observations and motion. The representation describes future end-effector targets relative to the current pose; it contains no target robot's joint-space action history. At deployment, the policy's relative targets are composed with the robot's current end-effector pose and expressed in the robot base frame. Inverse kinematics, motion planning, or a Cartesian controller then produces executable motion. My interpretation is that portability transfers the supervision problem into this interface: it reduces robot dependence during collection while making calibration, tracking, controller conventions, and contact compatibility decisive at deployment. A fluent human demonstration is therefore useful evidence of task structure, but it is not by itself a successful robot trial. e04-umi
Table 7, including data-source legend. Source membership expands across model generations, while effective proportions remain unresolved. Original paper, p. 36 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. Decode the five icons using the retained legend before reading a model row. The left half covers the earlier systems and the right half continues into 2026. Follow related model generations rather than treating every row as an independent experiment. In the pi series, the survey records robot data for pi0, robot plus general data for pi0.5, and an additional egocentric source for pi0.7. LingbotVA adds egocentric and general data to its earlier robot, UMI, and simulation recipe in version 2.0. These are source-presence observations; the icons contain no quantities, sampling weights, loss weights, or allocation across training stages. e08-recipese12-wame17-mixture-gap
What it supports. The table supports the authors' observation that heterogeneous recipes have become common. It also includes robot-only systems such as DreamZero. Combined with the adjacent text, this argues for testing each source's contribution instead of equating more icons with a better policy.
Where the evidence stops. The table labels VideoVLA as VLA, while Section 7.5.1 discusses it among WAM action-generation designs. These families are not cleanly disjoint here. Neither the labels nor the source icons establish comparative performance.
5.2 Separate action meaning from the coordinate system carrying it
Suppose two robots contribute equally long action vectors. The survey explains why that alone does not align their supervision: the same index may control different physical quantities. Embodiment-specific projectors retain native interfaces and share an internal representation. Zero padding standardizes tensor size without guaranteeing physical correspondence. Semantic slots assign common meanings and mask unavailable controls. Geometry is a separate decision. Robot-base coordinates suit conventional controllers, camera coordinates can align visually similar motions, and wrist coordinates separate local articulation from global hand movement. My deduction is that an aggregation pipeline needs two independently checked contracts: what each field means and how its reference frame is defined. The survey calls for documenting axes, handedness, units, tool-center point, rotation parameterization, and absolute versus delta commands, while explicitly leaving the best coordinate convention empirically unresolved. e09-alignmente16-transfer
5.3 Ask what a predicted future is allowed to support
A video may reveal that an object moved without revealing the robot action that caused it. The survey accordingly separates action-free temporal prediction from action-conditioned interaction. The former supplies broad dynamics priors; the latter grounds those priors for control. It then distinguishes three learned-simulator roles. Policy training uses imagined transitions to update behavior. Policy evaluation predicts progress, failures, or relative checkpoint quality without necessarily updating the policy. Synthetic-data generation creates videos whose action labels may still need an inverse-dynamics or latent-action model. My interpretation is that each role demands a different validation target: prediction accuracy, policy-ranking validity, and executable action consistency cannot substitute for one another. A visually convincing rollout could satisfy an appearance criterion while misleading all three downstream uses if the generator fails to preserve the effects of candidate actions. e12-wame13-learned-simulation
5.4 Training and inference
During training
These are reviewed training patterns, not a training procedure proposed by the survey. Embodied brains convert interaction into affordances, waypoints, and task structure. VLAs directly supervise discrete action tokens or continuous diffusion/flow-matching action heads. Video-derived latent actions and geometric motion proxies require additional robot grounding before they become controls. e10-braine11-vla
For WAMs, the survey describes action-free future prediction followed by action-conditioned adaptation. It distinguishes continuous action denoising from autoregressive token prediction: Motus uses interacting video/action experts, whereas WorldVLA jointly predicts discretized visual and action tokens. These examples illustrate architectural variety, not a universal training objective. No single optimizer, frozen-module schedule, compute budget, or fitted pyramid equation is prescribed. e12-wame17-mixture-gape22-survey
During inference
UMI deployment composes a predicted relative trajectory with the current end-effector pose, transforms targets into the robot base frame, and uses inverse kinematics, motion planning, or Cartesian control to execute them. Camera placement, tracking, contact geometry, and actuator dynamics can still break transfer. e04-umi
A learned world simulator may optimize a policy in imagined transitions, estimate a policy's performance, or generate synthetic videos with recovered pseudo-actions. These are distinct uses. Evaluation must preserve action effects and policy rankings; attractive video alone is insufficient. The survey defines no single deployed controller or inference schedule. e13-learned-simulatione22-survey
5.5 Implementation flow
- Locate the physical interaction loop
Real-robot trajectories record sensing, proprioception, actions, and actual outcomes on an embodiment. Scripted collection, teleoperation, and human intervention cover different state distributions. Intervention data specifically retain policy-induced errors and corrective behavior. Physical fidelity remains expensive, and data sharing still requires compatible sensors, coordinate frames, and controls. e03-robot
- Remove the robot from collection while retaining action structure
UMI combines portable manipulation interfaces with visual and inertial tracking and gripper state. Relative end-effector targets replace robot-specific joint commands. Human video offers broader everyday interactions but requires annotation, reconstruction, and retargeting to supply robot-oriented targets; sensor measurements, estimated poses, and generated action labels have different error sources. e04-umie05-human
- Separate synthetic interaction from general priors
Simulation generates controllable trajectories and privileged labels, but sensing, kinematics, and dynamics can differ from deployment. Replay expands demonstrated behaviors without guaranteeing new strategies. General data supplies semantics, spatial grounding, geometry, temporal memory, and planning priors. Neither broad semantic coverage nor plausible synthetic observations alone resolves physical execution. e06-simulatione07-general
- Align meaning as well as tensor shape
The survey distinguishes embodiment-specific projectors, fixed-size zero padding, and semantic action slots. Padding equalizes dimensions without ensuring identical meanings; slots give corresponding fields a physical interpretation and mask unavailable controls. Separately, robot-, camera-, and wrist-centric frames address geometry. Frame origin, axes, handedness, tool center, rotation representation, units, and absolute-versus-delta mode must be specified. e09-alignment
6. Experiments & results
This survey organizes manipulation data by what becomes easier to collect and what remains grounded in a robot's actions. Its five-layer pyramid connects real-robot, UMI, human-video, simulation, and general data to embodied reasoning, action generation, and world prediction. The useful output is a framework for choosing and aligning supervision; it does not establish an optimal mixture or introduce a new policy. The authors explicitly warn that more sources and more trajectories need not produce better control.
This survey proposes no new executable model architecture, shared policy benchmark, or controlled ablation. Its quantitative tables inventory datasets; Table 7 lists model/source membership; Figure 4 is a heuristic spatial diagnostic. The edition therefore uses the pyramid for mechanism, dataset statistics for quantitative context, and a diagnostic excerpt in the ablation section without treating it as an ablation experiment. The source itself identifies limited controlled evidence on data mixtures and coordinate conventions; no success-rate plot or optimal-mixture curve is fabricated. e22-surveye18-diagnostice09-alignmente17-mixture-gap
6.1 Read the original evidence
Table 1. Robot data grows along several dimensions that a trajectory count alone cannot describe. Original paper, p. 9 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. Read each row horizontally before comparing its trajectory count with another dataset. The slash in Traj. / Hours separates two quantities. Embod. counts distinct embodiment types; S, D, and H denote single-arm, dual-arm, and humanoid configurations. Calib. means calibrated camera extrinsics, Dex. marks dexterous-hand data, Anno. means within-episode subtask labels, and Mobile marks mobile manipulation. Locate RoboMIND and RoboMIND 2.0 to compare scale with tasks, modalities, and these flags. Dashes retain missing entries; they are not numerical zeros. The table inventories heterogeneous resources, so larger totals do not imply equal episode duration, task difficulty, or coverage quality. e19-robot-statisticse03-robote14-tactile
What it supports. RoboMIND 2.0 is listed with 310K trajectories, 1000+ hours, 739 tasks, and six embodiment types; RoboMIND lists 107K, 305.5 hours, 479 tasks, and four types. This supports a concrete statement about expanded inventory coverage, without providing a matched measurement of policy improvement.
Where the evidence stops. These are statistics compiled by the survey, not independently verified dataset releases or new experiments. Capability checkmarks indicate reported data properties, not accuracy, standardized sensing, successful transfer, or reliable recovery behavior.
Table 2. Portable collection scales demonstration volume while preserving differences in embodiment and contact sensing. Original paper, p. 15 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. Use the Arm Type and End Effector columns to establish what kind of motion a demonstration represents, then inspect Tasks and Demos together. UMI-style resources include single-arm and bimanual interfaces, parallel grippers and dexterous hands. The last column is especially easy to overread: DexUMI and ManipForce explicitly say Force / Torque, whereas other rows use tactile checkmarks or crosses. Keep those distinctions rather than collapsing every contact-related channel into one sensor modality. The accompanying deployment discussion explains why these robot-free demonstrations become useful: relative end-effector motion can be transformed and retargeted, although the collection device does not record a target robot's joint dynamics. e20-umi-statisticse04-umi
What it supports. The original UMI row lists 2.5K demonstrations across five tasks; FastUMI-100K lists 100K across 32. Both rows mark tactile sensing absent. The table therefore illustrates expansion of a collection resource while also making a missing interaction signal visible.
Where the evidence stops. Demonstrations are not robot evaluation trials. Different interfaces, tasks, tracking systems, and contact conditions prevent this inventory from isolating a hardware or scaling effect. Entries such as Diverse remain qualitative task descriptions, not recovered counts.
Table 5. Synthetic data volume and interaction coverage answer different questions. Original paper, p. 25 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. Begin with Tasks and Deform. before following the eye-catching scale column. Some entries concern grasp samples, while others concern task-level manipulation trajectories; Section 5.3 explicitly distinguishes these units of experience. Arm uses S for single-arm and D for dual-arm. Embod. describes embodiment count, with dexterity and mobility shown separately. Compare MimicGen and InternData-A1 as examples of different task and embodiment coverage, then inspect the missing-hour entries rather than attempting to convert every total to time. The table's Rigid, Soft, and Artic. labels describe supported interaction categories, not measurements of physical fidelity or transfer quality. e21-simulation-statisticse06-simulation
What it supports. The survey lists MimicGen with 50K trajectories, 16 tasks, and three embodiments, and InternData-A1 with 637K trajectories, 7.4K hours, 70 tasks, and four embodiments. The broader inventory shows why scale must be read alongside the type of supervision and interaction covered.
Where the evidence stops. No common success metric or evaluation protocol appears here. A grasp sample and a complete manipulation episode are not interchangeable, and synthetic coverage cannot establish real-world contact fidelity without separate evaluation.
6.2 Results and evaluation conditions
| Task & protocol | Reported result | Comparison & interpretation |
|---|---|---|
| Real-robot dataset scale and coverage inventory Survey Table 1; published dataset statistics, not a common train/test benchmark or policy evaluation. | RoboMIND 2.0: 310K trajectories, 1000+ hours, 739 tasks, 6 embodiment types. Trajectories; hours; tasks; embodiment types | RoboMIND: 107K trajectories, 305.5 hours, 479 tasks, 4 embodiment types. The inventory documents expanded scale and coverage. It supplies no controlled downstream success comparison, uncertainty estimate, or attribution of improvement to dataset size. e19-robot-statistics |
| UMI demonstration inventory Survey Table 2; robot-free collection resources with different task sets and interfaces; no shared evaluation split. | FastUMI-100K: 100K demonstrations across 32 tasks; single-arm gripper data; tactile marked absent. Demonstrations and task count | UMI: 2.5K demonstrations across 5 tasks; single/bimanual gripper data; tactile marked absent. Collection scales differ substantially, but these are dataset descriptions, not an ablation of collection hardware, tracking quality, or robot success. e20-umi-statistics |
6.3 Ablations and diagnostic examples
Read component removals and qualitative examples within their stated evaluation conditions.
Figure 4, lower-right plot and example frames (diagnostic excerpt). Spatial keyframes provide a limited diagnostic of repeated interaction geometry. Original paper, p. 7 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. Follow the colored connector lines from selected positions in the three-dimensional plot to the three frames on the right. These connectors associate positions with images; they are not action-flow arrows or a learned controller. The axes describe spatial coordinates in meters. The caption says interaction and grasp keyframes were selected using a heuristic following PerAct, with an analogous procedure for human hands. Read the dense regions as where the selected spatial samples accumulate. The survey's simulation discussion uses Figure 4 to caution that skill-constrained generation can repeat interaction regions despite a large dataset. It does not define point colors as success labels. e18-diagnostic
What it supports. This visualization makes spatial concentration inspectable: many selected keypoints occupy recurring regions, and the linked frames show the interaction context. It motivates checking behavioral diversity separately from raw trajectory volume. It does not measure how much additional robot competence those trajectories supply.
Where the evidence stops. The panel's title lettering is unavailable in the supplied render, so the excerpt is identified by its position and imagery. No color-to-outcome legend is given. This is a heuristic diagnostic, not a controlled ablation or generalization score.
7. Analysis & limitations
7.1 What the evidence leaves open
The pyramid is qualitative. Figure 4 uses heuristic interaction keyframes as a spatial-diversity proxy, without demonstrating that larger scatter predicts transfer. Different tasks, workspace bounds, and extraction choices prevent interpreting these panels as a controlled ranking. e02-pyramide18-diagnostic
Reported corpus volumes differ in modality, processing, filtering, and training stage. The authors explicitly leave source proportions and stage allocation unresolved because controlled, compute-matched comparisons are limited. Table 7 records mixtures, not their effectiveness. e08-recipese17-mixture-gap
Tactile data remains sensor-specific and poorly standardized; successful demonstrations underrepresent failure onset, causes, and recovery. Human hand inpainting can reduce visual mismatch while leaving kinematics and contact inconsistent. These gaps limit what even a large mixture can teach. e14-tactilee15-failuree16-transfer
7.2 Questions for discussion
- Which target capability justifies adding a particular data layer under a fixed collection and training budget?
- How should reconstructed human actions expose uncertainty about contact feasibility?
- Can a spatial keyframe-diversity diagnostic predict transfer after task and workspace differences are controlled?
8. Reproducibility audit
8.1 Requirements and known gaps
Reconstructing the survey's inventories would require versioned dataset definitions and consistent counting units. Preserve missing values and distinguish trajectories, hours, tasks, and grasp samples. Source-membership icons cannot recover sampling ratios, preprocessing, or stage-specific allocation; these would require the individual training reports. e19-robot-statisticse20-umi-statisticse21-simulation-statisticse08-recipese17-mixture-gap
A reader-proposed test should separate action-interface correctness from data-mixture benefit: first verify coordinate and semantic alignment on identical trajectories, then compare source mixtures under equal training compute and equal target-robot adaptation. Measure physical task completion separately from prediction quality. The illustrated checks specify controls; neither was run. e09-alignmente13-learned-simulatione17-mixture-gap
8.2 Proposed reproduction checks
The following checks are proposals motivated by the paper. They have not been run as part of this reading.
Check 1: Does a camera frame improve transfer after action semantics are fixed?
Reader-proposed, not performed: use the same task-matched trajectories from two robot embodiments, with one fixed semantic-slot interface and identical validity masks. Compare robot-base and calibrated camera-frame actions using the same backbone, data split, optimization budget, and target-robot adaptation. Before training, round-trip both encodings into the original controller frame and check translation, rotation, and gripper commands against the recorded targets. Include a deliberately perturbed extrinsic calibration as a sensitivity control. Evaluate held-out scenes and physical task completion with repeated seeds and trials. If the aligned camera representation does not improve transfer, the result would constrain a proposed camera-frame advantage rather than refute the need for correct coordinate metadata. e09-alignmente16-transfer
Check 2: Does human-video diversity help control beyond future prediction?
Reader-proposed, not performed: hold a world-model architecture and action-free future-prediction objective fixed. Compare pretraining on robot videos alone, repeated to fill the budget, against a predeclared robot/human-video mixture under matched processed tokens and measured training compute. Keep resolution, temporal sampling, subsequent action-conditioned robot training, and target-robot adaptation identical. Separate training scenes and objects from evaluation, and report actual source exposure in each arm. Evaluate both future-prediction quality and real-robot task completion, including perturbed starts that can require recovery. Better prediction without better execution would show that the added observational diversity has not established a control benefit under this recipe. Report uncertainty across seeds and trials rather than selecting one favorable run. e08-recipese12-wame15-failuree17-mixture-gap
8.3 Reading coverage
Visual audit: The title and author page, conceptual pyramid, trajectory diagnostic, collection diagrams, four retained inventory tables, and all declared pages supporting method, training, evaluation boundaries, and proposed checks were visually inspected. Every final crop was viewed, including the corrected diagnostic bounds and complete source-icon legend. Figure 4 connector lines were checked as image associations, and the UMI-to-robot direction was checked against Figure 5's caption and Section 3.3. Some lettering elsewhere in Figures 1, 2, 4, 5, and 6 is unavailable in the supplied renders; no missing label was drawn or inferred as a numerical result. Figure 1 is cropped to its readable pyramid; Figure 4 is explicitly a positional excerpt. Figures 3, 7, and 8 and Tables 3, 4, and 6 were text-read but not visually inspected. The reference list was completely text-read; its remaining pages were not visually inspected. No appendix is present and separate supplements were not verified.
PDF pages inspected for this edition: 1, 3, 4, 7, 8, 9, 12, 13, 15, 16, 18, 21, 22, 25, 26, 27, 28, 29, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46. Appendix coverage: not present.
Original figures and tables remain the work of the source’s authors. Extractions preserve their scientific content; any HTML wrapper layout is disclosed with each figure. The surrounding reading notes are our own.
Text reading scope & known omissions
- Title, author credits, abstract, and contents, pp. 1–2
- 1 Introduction, pp. 3–7
- 2 Real-Robot Data: all subsections, pp. 7–14
- 3 UMI Data: all subsections, pp. 14–16
- 4 Egocentric and Exocentric Data: all subsections, pp. 16–22
- 5 Simulation Data: all subsections, pp. 22–29
- 6 General Data: all subsections, including 3D and physical/causal/failure reasoning, pp. 29–35
- 7 Data Applications in Embodied Foundation Models: all subsections, pp. 35–43
- 8 Challenges and Future Directions: all six subsections, pp. 43–45
- 9 Conclusion, pp. 45–46
- References [1]–[453], pp. 46–72; all 28 supplied text chunks read individually
Outside the original text pass
- Text extraction does not reconstruct figure images; inspect the retained PDF for figures and equation/table layout.
- Separate supplemental material availability has not been fully verified.
- The supplied title and all 29 author names match the catalog. The inspected artifact is arXiv:2607.24744v2, with an 8 August 2026 margin date and an August 11, 2026 title-block date. The catalog gives August 8, 2026. These dates are preserved; no earlier version or revision history was supplied, so differences from v1 were not compared.
- Visual inspection supplements the extracted text for the pages declared in the edition. Some lettering in the overview and collection figures is blank in the supplied PDF renders; missing lettering was not reconstructed. The retained Figure 1 excerpt and Figure 4 diagnostic preserve readable scientific content.
- Figures 3, 7, and 8 and Tables 3, 4, and 6 were read through the supplied text but were not visually inspected. No quantitative claims depend on their unread visual details.
- The cited datasets and individual model papers were not independently opened. The linked curation repository and project were not inspected, code was not executed, and no experiments were reproduced.
The visual audit above records the subsequent illustrated pass.
8.4 Traceable evidence
e01-identityPDF p. 1, title, complete author/affiliation block, date, and arXiv margin
The title and 29 authors match the supplied catalog. The artifact states arXiv:2607.24744v2, 8 Aug 2026, while the title block says August 11, 2026; affiliations are supplied as institutional abbreviations.
Go to primary source ↓e02-pyramidPDF p. 1, Figure 1 pyramid; pp. 3–4, Section 1, organizing principles and layer ordering
Five categories are ordered from real-robot through UMI, ego/exo, simulation, and general data; scalability and robot alignment are primary, with four complementary dimensions. The text explicitly rejects monotonicity along every property.
Go to primary source ↓e03-robotPDF p. 8, Figure 5 caption; p. 12, Section 2.3.3; p. 13, Sections 2.4–2.5
Real-robot collection includes scripts, teleoperation, and intervention; interventions add corrective actions at policy-induced states. Physical grounding comes with collection cost and cross-platform heterogeneity.
Go to primary source ↓e04-umiPDF p. 8, Figure 5 caption and UMI-to-robot arrow; p. 15, Sections 3.2.2–3.3; p. 16, Section 3.4
Visual-inertial tracking yields relative gripper trajectories without robot joint states. Deployment composes relative targets with current end-effector pose and converts them to robot controls; tracking, calibration, and embodiment gaps remain.
Go to primary source ↓e05-humanPDF p. 16, Section 4.1; p. 21, Section 4.3.4; p. 22, Section 4.4
Human recordings gain semantic and geometric supervision through sensing and post-processing. Robot-compatible targets require coordinate transformation, IK, retargeting, and filtering; geometric plausibility does not guarantee feasible contact.
Go to primary source ↓e06-simulationPDF pp. 26–27, Sections 5.3–5.4; pp. 28–29, Sections 5.6–5.7
Simulation supplies privileged labels and repeatable trajectories; demonstration expansion reuses seed behavior and is limited when qualitatively new strategies are required. Observation, kinematic, and dynamic mismatch constrain transfer despite scalable generation.
Go to primary source ↓e07-generalPDF p. 35, Section 6.9 and preceding grasp-resource distinction
General data contributes complementary semantic, spatial, geometric, temporal, planning, and physical-reasoning priors with weak action grounding; generated annotations require verification.
Go to primary source ↓e08-recipesPDF pp. 35–37, Section 7.2.1; p. 36, Table 7, legend and pi-series, LingbotVA, DreamZero rows
Mixtures diversify across model generations; LingbotVA expands from robot/UMI/simulation to all five sources. Robot-only systems remain viable. Reported hours are not directly comparable across processing and training stages.
Go to primary source ↓e09-alignmentPDF pp. 37–38, Section 7.2.2, structural and geometric action representations
Projectors, zero padding, and semantic slots solve different structural problems. Robot-, camera-, and wrist-centric frames require explicit geometry and control conventions; controlled frame comparisons remain limited.
Go to primary source ↓e10-brainPDF pp. 38–39, Sections 7.3–7.3.2
Embodied brains use action-free understanding tasks and transform interaction data into affordances, trajectories, and task structure, rather than necessarily imitating raw controls.
Go to primary source ↓e11-vlaPDF pp. 40–41, Sections 7.4.1–7.4.2
VLAs learn action tokens or continuous diffusion/flow heads; video-derived latent actions and geometric proxies require robot-specific grounding. High-level intermediate supervision can also guide action generation.
Go to primary source ↓e12-wamPDF pp. 41–43, Section 7.5, especially p. 42, action-labeled and action-free training
The survey separates observational future prediction from action-conditioned dynamics and continuous denoising from autoregressive action prediction. Motus has video/action experts; WorldVLA predicts discretized visual/action tokens.
Go to primary source ↓e13-learned-simulationPDF p. 28, Sections 5.5.1–5.5.3 and concluding paragraph
Learned simulators support policy optimization, evaluation, and synthetic data generation. Evaluation needs controllable action effects and valid relative policy rankings; pseudo-action recovery and hallucinated dynamics introduce errors.
Go to primary source ↓e14-tactilePDF p. 43, Section 8.1
Touch exposes contact information missing from vision but is not a standard dataset modality; hardware dependence, formats, and coverage impede scale and reuse.
Go to primary source ↓e15-failurePDF pp. 43–44, Section 8.2
Success-focused curation omits error states. Richer failure data should retain preceding context, onset, causes, state changes, corrective actions, and recovery outcomes.
Go to primary source ↓e16-transferPDF pp. 44–45, Sections 8.3–8.5
Scalable sensing requires calibration and synchronization; common storage does not ensure action semantics. Human-to-robot transfer includes morphology and contact gaps that visual inpainting cannot resolve.
Go to primary source ↓e17-mixture-gapPDF p. 45, Section 8.6, final paragraph
Optimal proportions and stage-wise allocation are unestablished; compute-matched ablations remain limited. The authors call for architecture-aware, stage-dependent comparisons of fixed, curriculum, and adaptive recipes.
Go to primary source ↓e18-diagnosticPDF p. 7, Figure 4 and caption, lower-right diagnostic excerpt; p. 29, Section 5.7, InternData-A1 discussion
Heuristic grasp/interaction keyframes visualize spatial coverage. The source uses InternData-A1 to illustrate concentrated action keypoints despite scalable skill-constrained generation; it does not report a policy ablation or diversity score.
Go to primary source ↓e19-robot-statisticsPDF p. 9, Table 1, RoboMIND and RoboMIND 2.0 rows; header and caption definitions
RoboMIND lists 107K trajectories, 305.5 hours, 479 tasks, four embodiments; RoboMIND 2.0 lists 310K, 1000+, 739, six. The latter marks calibrated cameras, dexterity, subtask annotations, and mobile manipulation.
Go to primary source ↓e20-umi-statisticsPDF p. 15, Table 2, UMI, FastUMI-100K, DexUMI, ManipForce, and modality columns
UMI lists 2.5K demonstrations/five tasks; FastUMI-100K lists 100K/32, both without tactile. DexUMI and ManipForce explicitly list Force / Torque in the tactile column.
Go to primary source ↓e21-simulation-statisticsPDF p. 25, Table 5, MimicGen, InternData-A1, DexGraspNet 2.0, and capability columns; p. 26, Section 5.3
MimicGen lists 50K trajectories, 16 tasks, three embodiments; InternData-A1 lists 637K, 7.4K hours, 70 tasks, four embodiments. Section 5.3 distinguishes task trajectories from grasp/dexterous samples in this inventory.
Go to primary source ↓e22-surveyPDF p. 1, abstract; pp. 45–46, Section 9 Conclusion
The paper contributes a data-centric synthesis, inventories, applications, and open challenges, rather than a newly trained policy with its own benchmark protocol.
Go to primary source ↓8.5 Primary sources
Data Pyramid for Embodied Manipulation: A Survey ↗
PDF · 43,864 extracted words
Source fingerprint
c4c64dc37b4b7aed9bea7f2499cd2b0110260bf07609c4fd458206b9bf410f82