PAPER REPORTENAll readings ↗

ManipArena: A Controlled Benchmark for Diagnosing Generalization in Real-Robot Manipulation

English reading report: Method, equations, original figures, experiments and reproducibility.

Authors: Yu Sun; Meng Cao; Yang Ping; Kaidong Zhang; Qingxuan Chen; Rongtao Xu; Liangwang Ruan; Xuecheng Chen; Dongxiu Liu; Yunxiao Yan; Zunnan Xu; Runze Xu; Charles Yang; Peilun Zhang; Xiaofan Li; Ruyi Gan; Liang Ma; Yuehao Yin; Jincheng Yu; Lufang Chen; Yuxin Liang; Peng Zhai; Hao Wang; Ivan Laptev; Ian Reid; Qian Wang; Xiaodan Liang

Affiliations: Sun Yat-sen University; X Square Robot; MBZUAI; Tsinghua University; University of Zurich

Source: 2603.28545 ↗ · Catalog record

Reading: 230 / 558 · 6 original figures & tables · ~20 min ·

1. Paper overview

In one sentence: Controlled physical trials expose how demonstration selection, language annotations and pretraining provenance change manipulation rankings, while limited coverage and inconsistent reporting constrain the conclusions. e02e03e04e07e08e10e16

At a glanceWhat to know
Research problem
Source description

Real-robot comparisons confound policy capability with hardware, resets, objects and scoring. Simulator evaluations remove some confounds but miss physical noise and contact effects. ManipArena standardizes physical conditions while varying declared task dimensions, so failures can be examined beyond a single success label. e02

Core mechanism
Source description

A benchmark spanning 15 tabletop and five mobile tasks contains 10,812 demonstrations, 13.5M frames and approximately 188 robot hours. Fixed lighting, cameras and workspace geometry make the physical setup part of the measurement protocol. e02

A key reported resultDemonstration-count scaling: 200/task: 312/500 and 42% SR.

Five-task score /500; mean SR. Unified π0.5; five representative tabletop tasks; fixed language/interface; 60k steps except the full-data 120k control.

Full data at 60k: 277 and 34%; full data at 120k: 248 and 26%. The selected subset wins in this sweep. Sampling and composition remain relevant; this does not establish a universal 200-demonstration optimum or an isolated causal benefit of balancing. e07

Reading caution
Reader analysis

Ten physical trials per task give limited repeatability evidence, especially the two-trial held-out band. No uncertainty intervals accompany the reported comparisons. Controlled booths, fixed embodiments and the common L1 interface restrict deployment and prompt generalization. e03e13

Core contributions

  • Source description

    A benchmark spanning 15 tabletop and five mobile tasks contains 10,812 demonstrations, 13.5M frames and approximately 188 robot hours. Fixed lighting, cameras and workspace geometry make the physical setup part of the measurement protocol. e02

  • Source description

    Task schemas, three language annotation levels, stratified trials and task-specific partial credit support diagnostic comparisons. Paired simulation contributes a limited data-source study with real-robot endpoints. e03e04e05e11

Figure 2. The controlled apparatus is part of the benchmark's method. Original paper, p. 5 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Start with the left photograph: the arrows label the green screen, studio lights, head camera, follower arms with cameras, and master arms used in the collection setup. The right photograph adds a height-adjustable column and a Mecanum-wheel omnidirectional base inside the labeled 3 m × 3 m enclosure. These arrows identify physical components; they are not neural information-flow arrows. Section 3.1 specifies fixed lighting, camera placement and workspace geometry, with both platforms operating at 20 Hz. Figure 2 therefore explains what is held stable when the authors vary demonstrations, language or fine-tuning regime. Algorithm 1 separately identifies teleoperation as the demonstration-collection step. e02e05e12e13

What it supports. The benchmark couples a fixed-base bimanual suite with a mobile suite that requires navigation-conditioned manipulation. Their respective 15 and five tasks share a controlled evaluation philosophy. The photographs support the apparatus description; capability measurements come from executed trials and the task rubrics, not from the photographs themselves.

Where the evidence stops. Fixed physical conditions improve comparability within this setup. They do not test arbitrary homes or factories. The five-task mobile results cover only one task-specific π0.5 configuration, so the two photographs do not imply equally broad model evaluation on both platforms.

2. Motivation

2.1 The problem and the proposed response

Source description

Real-robot comparisons confound policy capability with hardware, resets, objects and scoring. Simulator evaluations remove some confounds but miss physical noise and contact effects. ManipArena standardizes physical conditions while varying declared task dimensions, so failures can be examined beyond a single success label. e02

2.2 What this reading follows

ManipArena asks how to make a robot evaluation diagnostically useful. It organizes demonstrations and tests around declared task schemas, runs policies in controlled physical booths, and awards points for intermediate subgoals as well as measuring near-completion. Read the figures as parts of an experimental instrument: the apparatus fixes physical conditions, the tables compare training interventions, and the OOD audit specifies what actually changed at test time. The strongest reported gains often come from recipe choices within one policy family. The source also contains conflicting model counts and per-task entries, so this edition preserves reported aggregates while identifying where independent reconciliation is needed. e02e03e04e07e08e10e16

3. Research context

We place the paper in the collection through its world–action interface. The catalog labels and the reading’s assessment are shown separately.

Catalog dimensionRecorded classification
Major categoryNot assigned
ArchitectureNot assigned
Prediction paradigmNot assigned
QuadrantNot assigned

This table preserves the labels recorded at reading time. The current major category is Benchmarks & simulators. View the current classification.

3.1 Evidence-based assessment

Classification assessment not applicable

Reader analysis

ManipArena is an evaluation benchmark encompassing VLA and WAM baselines, not itself a world-action architecture. One Model versus Hybrid and joint future/action prediction versus inverse dynamics are therefore not applicable to the benchmark. The paper's family labels do not supply enough internal architecture evidence to classify every baseline. e02e06e10

This is the collection’s architectural analysis, not a new related-work survey. Benchmark comparisons and their protocols appear in Section 6.

4. Problem formulation

4.1 Inputs and outputs

InputsOutputs
  • Task schemas and teleoperated demonstrations
  • Head/wrist visual observations, platform-specific proprioception and language instructions
  • Released 56D tabletop or 62D mobile state/action records; model-specific adapters select their interfaces
  • Executed robot trajectories
  • Per-trial subgoal scores, task/category totals and thresholded success rates

4.2 Equations and their role

Si(t)=k=1Kiwi,k1[gi,k achieved on trial t],k=1Kiwi,k=10S_i^{(t)}=\sum_{k=1}^{K_i}w_{i,k}\,\mathbf{1}[g_{i,k}\text{ achieved on trial }t],\qquad\sum_{k=1}^{K_i}w_{i,k}=10
Equation (3): task i has K_i rubric subgoals g_{i,k}, weights w_{i,k}, and trial score S_i^{(t)}. The indicator counts achieved subgoals. Consult the task rubric for exceptions such as wrong-fruit penalties. e04
S(I)=iIt=110Si(t),SR(I)=1IiI110t=1101[Si(t)9]S(\mathcal I)=\sum_{i\in\mathcal I}\sum_{t=1}^{10}S_i^{(t)},\qquad \mathrm{SR}(\mathcal I)=\frac{1}{|\mathcal I|}\sum_{i\in\mathcal I}\frac{1}{10}\sum_{t=1}^{10}\mathbf{1}[S_i^{(t)}\geq9]
Equation (4): I is the evaluated task subset. S sums progress points; SR averages the fraction of near-complete trials across tasks. Tables express SR as a percentage, not as a normalized partial-credit score. e04

5. Method in detail

5.1 Follow one trial from schema to score

Reader analysis

The experimental unit is a task configuration, not simply a task name. A schema selects object properties, layout and distractors; data collection resets the scene, records teleoperation and attaches three language annotations. At evaluation, the same schema defines in-domain, shifted and held-out bands, after which the policy actually executes. The rubric then converts observable subgoals into progress points. For put_ring, moving the rod to the center, picking the ring and threading it each earn three points, while retraction earns one. Equation (4) separately asks whether the trial reached nine points. Reader interpretation: this separation helps distinguish useful preparation from near-completion, but the rubric remains an operational definition of performance. The summary table does not provide every mobile stage weight, so those cases need the complete scoring specification before replication. e03e04e05e17

Table 15. The final trial band tests different kinds of novelty across tasks. Original paper, p. 20 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Read the two large blocks before interpreting any held-out score. Object OOD means the final two trials introduce previously unseen objects, such as sunglasses in put_glasses or a black plastic spoon in put_spoon. Configuration OOD instead changes arrangements, orders or procedures while reusing in-distribution objects: press_button holds out button orders, and put_ring uses rare ring/rod-position pairings. The middle column records the earlier T5–T8 shift; arrows there enumerate appearance changes rather than transitions executed by a policy. The right column specifies T9–T10. Trials T1–T4, defined in Section 3.3, sample in-domain configurations and are not listed in this table. e03e05e13e15

What it supports. Only seven of the fifteen tabletop tasks use object OOD in their final trial band. Eight test configuration OOD. An aggregate tabletop score therefore combines different generalization demands; the word held-out alone does not establish object novelty, semantic novelty or robustness to unrestricted deployment environments.

Where the evidence stops. This table specifies the intended audit conditions, not per-band performance or proof of zero overlap with upstream pretraining. Each task has only two final-band trials in the stated protocol. The full schema assets and sampled configurations would be needed for a deeper overlap audit.

5.2 Understand why training context can help an L1-only policy

Reader analysis

The language experiment varies what accompanies a demonstration during fine-tuning. L2 adds object and spatial details to the goal; L3 adds scene description. At evaluation, neither model receives its richer training annotation: the shared prompt is L1. This distinction matters because the reported gains are not explained by simply giving one tested policy a more informative command. The aggregates favor L2 for execution score and L3 for semantic score and overall success rate. Reader interpretation: language may shape what visual and action distinctions training emphasizes, even when deployment instructions are abstract. That is a plausible explanation, not an identified internal mechanism. A faithful replication must also reconcile Table 9's task entries with its aggregate row before analyzing which tasks account for the reported trade-off. e06e08e16

5.3 Ask what a controlled comparison actually holds fixed

Reader analysis

Matched robot trials control the measurement interface, but they do not equalize every cause of policy performance. The source keeps implementation and evaluation conventions fixed within a model family while comparing specialists, unified policies, data quantities or language fields. Across families, upstream pretraining still differs. The opposite effects of task sharing therefore support an interaction between model provenance and fine-tuning regime, while leaving the authors' pretraining explanation correlational. The simulation diagnostic illustrates the same discipline from another direction: changing training data is evaluated through the unchanged real-robot endpoint, and its benefit reverses across tasks. Reader interpretation: ManipArena is most useful as a framework for asking narrower, falsifiable questions. A single aggregate ranking cannot establish that one architecture reasons better, generalizes to new objects everywhere, or will transfer to unrestricted environments. e06e09e10e11e13e15

5.4 Training and inference

During training

Source description

Baselines are fine-tuned as task-specific specialists or unified policies. Within each controlled study, the authors hold implementation, action representation, normalization, evaluation prompts and scoring fixed except for the intended factor. Upstream pretraining differs. Scaling varies demonstration count; language studies vary training annotations while testing with L1. e06e07e08

Reader analysis

This benchmark does not specify a new neural architecture or common learned objective. The supplied PDF does not provide complete optimizer, learning-rate, batch-size, frozen-module or checkpoint-revision recipes for reconstructing every baseline. e06e14e17

During inference

Source description

Policies receive the shared high-level L1 instruction through their model-specific observation/action interfaces and execute on the robot. The benchmark measures the resulting trajectory; it does not prescribe an imagined-future planner or unified future/action decoder. Simulation-trained checkpoints are also evaluated physically. e03e06e11

5.5 Implementation flow

  1. Construct controlled task variation

    A schema declares object, layout, distractor and held-out conditions. Sample a configuration, reset the scene, collect a teleoperated trajectory and generate its L1/L2/L3 annotations. Some dimensions are jointly sampled or refreshed at batch boundaries; independent sampling is only the common case. e05

  2. Expose the physical interface

    Both platforms operate at 20 Hz. Tabletop uses a head camera and two wrist cameras; mobile adds a lifting column and omnidirectional base. LeRobot v2.1 records include end-effector fields, joint positions, velocities and currents, plus mobile extras. Released channels do not establish that every baseline consumes all of them. e02e05

  3. Execute stratified trials

    For each tabletop task, trials 1–4 are in-domain, 5–8 introduce visual/configuration shifts, and 9–10 use the strongest held-out condition. Seven tasks use object OOD; eight reuse objects with held-out arrangements, orders or procedures. Reset and execute the policy before scoring. e03e15

  4. Separate progress from completion

    Each trial has ten rubric points; success means at least nine. Fifteen tabletop tasks therefore yield a maximum 1,500 points. A high partial-credit total can coexist with low success. The fruit task additionally penalizes wrong picks, and several mobile rubric summaries omit individual weights. e04e17

6. Experiments & results

ManipArena evaluates manipulation policies through controlled physical robot trials, task schemas and subgoal scoring. Its strongest lesson is that training recipes and model provenance affect rankings alongside architecture. Language grounding and demonstration selection produce substantial reported gains, but small trial counts, restricted environments and internal reporting inconsistencies limit causal and generalization claims.

6.1 Read the original evidence

Table 5. Model-family rankings depend on fine-tuning regime and provenance. Original paper, p. 9 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Read each Sc/SR pair together. Execution contributes at most 1000 points and semantic tasks contribute 500, so their raw category totals do not have equal scale. Spec. means separate task-specific specialists; Unified means one multi-task setting. The π0.5 and WALL-OSS groups each contain both regimes, whereas X-VLA, Fast-WAM, DreamZero and Wall-WM each have a unified column. The full crop deliberately preserves the final Wall-WM column even though Section 4.4 and Figure 5 describe a seven-configuration landscape without that column. All scores here are tabletop aggregates under the common ten-trial protocol and L1 instruction interface. e03e06e09e10e16

What it supports. WALL-OSS-0.5 specialist reports 951/1500 and 38% SR, the highest total score shown. π0.5 unified reports 703 and 20%; WALL-OSS unified reports 865 and 34%. The within-family direction differs: task sharing helps the reported π0.5 total but reduces the WALL-OSS total. This is evidence about evaluated configurations.

Where the evidence stops. Eight displayed configurations conflict with the prose count of seven. Printed rounded cells also require care when comparing summaries. The authors' embodiment-aligned pretraining explanation is correlational; neither matching robot trials nor a leaderboard ranking isolates architecture from upstream data.

Table 11. Simulation is tested as a training-data source with physical endpoints. Original paper, p. 17 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. The column headers identify how the checkpoint was trained. They do not identify three evaluation environments: all columns were evaluated on the real robot. Sc is a task's partial-credit total out of 100, and SR is success percentage. Follow each task horizontally before comparing tasks vertically. put_blocks changes little, press_button improves substantially with simulated training, and pick_fruits favors real training. Cotrain 50% is the source's mixed real/sim recipe. The table is deliberately small because the paired simulation extension covers only these three tasks. Its role is to screen data-source choices for physical validation, as Section 4.5 explicitly states. e04e11e13e17

What it supports. For press_button, simulated-data training reports 65 points and 50% SR, versus 31 points and 0% for real data. Yet pick_fruits scores 78 with real training and 51 with simulated training. The mixed recipe scores 40 and 68 on those tasks, so it does not uniformly dominate either single-source recipe.

Where the evidence stops. These task-dependent results do not establish general simulator fidelity or a universal mixing ratio. The PDF gives no uncertainty intervals here and no complete simulation-generation recipe. Additional tasks and repeated physical trials are needed to assess transfer reliability.

6.2 Results and evaluation conditions

Task & protocolReported resultComparison & interpretation
Tabletop model comparison

15 physical tabletop tasks; ten stratified trials each; common L1 prompts; specialist and unified fine-tuning.

WALL-OSS-0.5 specialist: 951/1500 and 38% SR in Table 5; Section 4.4 reports 950.8/1500.

Partial-credit score /1500; mean success rate

Table 5: π0.5 unified 703/1500, 20% SR; WALL-OSS unified 865, 34%; X-VLA 442, 7%; Fast-WAM 465, 11%; DreamZero 500, 11%; Wall-WM 868, 31.33%.

The best reported score is approximately 63.4% of available points, not 63.4% success. Architecture and pretraining are confounded; the additional Wall-WM column contradicts the narrative's seven-configuration count. e10e16

Demonstration-count scaling

Unified π0.5; five representative tabletop tasks; fixed language/interface; 60k steps except the full-data 120k control.

200/task: 312/500 and 42% SR.

Five-task score /500; mean SR

Full data at 60k: 277 and 34%; full data at 120k: 248 and 26%.

The selected subset wins in this sweep. Sampling and composition remain relevant; this does not establish a universal 200-demonstration optimum or an isolated causal benefit of balancing. e07

Training-language granularity

Unified π0.5 on 15 tabletop tasks; training L1, L2 or L3; all evaluations use L1.

Table 9 total row: L2 867.25, 32.7% SR; L3 801.75, 35.3% SR.

Reported total score /1500 and mean SR

L1 total row: 702.5, 20.0%. Table 4 reports rounded execution scores 421/535/432 and semantic scores 282/333/370 for L1/L2/L3.

Reported aggregates favor L2 for execution and L3 for semantics and success. They are transcribed as printed; Table 9 rows and aggregate summaries do not fully reconcile. e08e16

Specialist versus unified fine-tuning

Same 15 tabletop tasks and evaluation protocol; comparison within each model family.

π0.5: +76.3; WALL-OSS-0.5: −86.0.

Unified minus specialist partial-credit points

Execution/semantic changes: +47.3/+29.0 for π0.5 and −59.0/−27.0 for WALL-OSS.

Task sharing has opposite reported effects. Embodiment-aligned pretraining is the authors' correlational explanation, not a controlled pretraining intervention. e09

Real-robot transfer from simulation

π0.5 checkpoints trained on real, simulated or 50% mixed real/sim demonstrations; three paired tasks; physical evaluation.

press_button: simulated 65/100, 50% SR.

Per-task score /100; SR

Real 31, 0%; mixed 40, 20%. On pick_fruits, real/sim/mixed scores are 78/51/68 and SRs 50%/30%/50%.

Simulation's value is task-dependent. The mixed checkpoint does not uniformly dominate, and simulator success alone is not the endpoint. e11

Scoped mobile manipulation

Only task-specific π0.5 on five navigation-conditioned tasks; ten trials per task.

put_clothes 4.5; hang_picture 0; organize_shoes 8; put_bottle 14; set_tableware 10; all 0% SR.

Task score /100; SR

Other mobile model entries are unavailable, shown as dashes.

These outcomes characterize one scoped configuration, not comparative mobile capability across model families. e12

6.3 Ablations and diagnostic examples

Read component removals and qualitative examples within their stated evaluation conditions.

Figure 3. More demonstrations do not monotonically improve the tested π0.5 recipe. Original paper, p. 7 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Read the table first: Data specifies demonstrations per task, Steps specifies the training budget, Sc. is total partial credit out of 500, and SR is mean success rate in percent. The plot repeats the score column for the same five-task evaluation. The purple point is 200/task; the final diamond and dagger-marked Full row use 120k steps rather than the 60k used elsewhere. The five tasks concern glasses, blocks, wire insertion, button pressing and fruit selection. Compare recipes within this figure before consulting the separate 15-task landscape: the two aggregations have different denominators and must not be interchanged. e07e13

What it supports. The 200/task checkpoint reports 312/500 and 42% SR, compared with 277 and 34% for full data at 60k steps. Doubling the full-data budget gives 248 and 26%, so that control does not recover the subset's reported advantage. The observation motivates attention to demonstration composition alongside count.

Where the evidence stops. The curve identifies a winner in this sweep, not a universal optimum. No error bars are shown. Subset composition and training variability remain unresolved, and the study does not independently establish that balancing alone caused the gain.

Table 4. Training-language detail changes the reported execution–semantic trade-off. Original paper, p. 8 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Columns describe training annotations, not different test prompts: every policy is evaluated using L1. L1 gives the task goal; L2 adds episode-specific object and spatial grounding; L3 adds a generated scene description to the grounded instruction. Read execution and semantic scores against their own maxima, 1000 and 500, before looking at Total /1500. Sc is partial credit and SR is success percentage, with success requiring at least nine points in a trial. The blue L2 cells emphasize stronger execution, while the pink L3 cells emphasize semantic performance. This table rounds totals and SR; Appendix F supplies more precise printed aggregate values. e04e08e16

What it supports. The reported execution score rises from 421 under L1 to 535 under L2, whereas the largest semantic score is L3's 370. L2 has the highest total score, 867, but L3 has the highest rounded overall SR, 35%. Progress and thresholded completion therefore favor different settings.

Where the evidence stops. Treat these as printed aggregates. Appendix Table 9 contains unreconciled per-task entries, including L1 arrange_cup; Table 4's L1 execution SR also differs from Table 5. Better scores do not by themselves identify an internal semantic-reasoning mechanism.

7. Analysis & limitations

7.1 What the evidence leaves open

Reader analysis

Ten physical trials per task give limited repeatability evidence, especially the two-trial held-out band. No uncertainty intervals accompany the reported comparisons. Controlled booths, fixed embodiments and the common L1 interface restrict deployment and prompt generalization. e03e13

Reader analysis

Internal inconsistencies prevent treating all printed cells as a reconciled dataset: Tables 5/8 contain eight configurations, whereas the prose says seven. Table 9 gives L1 arrange_cup as 18, versus 47 for unified π0.5 in Table 8; its total row does not equal the sum of its printed L1 task scores. Table 4 gives L1 execution SR 15%, versus 16% in Table 5. These are preserved, not repaired. e16

7.2 Questions for discussion

  1. Would the scaling advantage persist under repeated, schema-stratified subsets and multiple training seeds?
  2. Would the language trade-off persist after reconciling per-task scores and separating object OOD from configuration OOD?

8. Reproducibility audit

8.1 Requirements and known gaps

Reader analysis

Reproduction requires the actual schemas, split membership, sampled trajectory IDs, calibration/reset procedures, model adapters, normalization assets and complete scoring rules. The PDF refers to released schemas but does not enumerate their full distributions. Several mobile rubric weights and detailed baseline optimization settings remain unspecified here. e05e06e17

Source description

Table 7 estimates unified π0.5 fine-tuning at 60k steps on eight A800 GPUs for 20 hours: 160 GPU-hours per checkpoint. A specialist uses 30k steps, eight A800s and six hours: 48 GPU-hours. These exclude upstream pretraining and physical reset/scoring labor; they are approximate accounting, not hardware-normalized efficiency measurements. e14

8.2 Proposed reproduction checks

The following checks are proposals motivated by the paper. They have not been run as part of this reading.

Check 1: Separate subset composition from demonstration count

Reader-proposed check: compare unified π0.5 trained on 200/task and full data at the same 60k-step budget, then repeat with independently sampled 200/task subsets and multiple training seeds. Hold language field, normalization, adapters and test prompts fixed; log schema coverage and exact trajectory IDs. Evaluate matched physical resets in randomized checkpoint order, reporting per-task and per-band scores, SR and uncertainty. Include a schema-stratified subset alongside an unstratified subset at the same size. If the advantage disappears across resampling or survives equally without stratification, a specific balancing explanation is weakened. A persistent gain tied to measured coverage would support a more precise sampling hypothesis. e05e06e07e13e15

Check 2: Reconcile language results, then test the grounding explanation

Reader-proposed check: first reconstruct all language-study aggregates from per-trial scoring records and resolve the L1 arrange_cup discrepancy without silently selecting one published value. Then train matched L1, L2 and L3 variants plus a control that swaps episode-specific grounding between demonstrations of the same task. Keep demonstration IDs, optimizer recipe, seeds and L1 test prompts matched; score execution and semantic subsets separately under blinded rubric review. If correct grounding reliably outperforms swapped grounding, that supports an information-alignment explanation. If the reported L2/L3 trade-off vanishes after reconciliation or repeated trials, it should not be treated as a stable training-design rule. e04e06e08e13e16

8.3 Reading coverage

Visual audit: The title/byline page, original Figures 1–5, Tables 1–15, scoring equations, both algorithms, implementation/resource tables, dataset layout, rubrics, limitations and OOD assignments were visually inspected. All six final crops were separately inspected; they preserve original pixels, labels and table entries. Figure 2 serves as the benchmark's physical-method visual because the paper introduces an evaluation instrument rather than a neural architecture. Every PDF page supporting retained method, numerical, training, evaluation and reproduction claims is included above. Pages 4 and 10–12 were read in the complete text but were not part of the image pass. Separate supplementary files and external schema/code assets were not inspected. The title-revision chain beyond the visible PDF rests on supplied acquisition evidence; no fresh web inspection is claimed.

PDF pages inspected for this edition: 1, 2, 3, 5, 6, 7, 8, 9, 13, 14, 15, 16, 17, 18, 19, 20. Appendix coverage: reviewed.

Original figures and tables remain the work of the source’s authors. Extractions preserve their scientific content; any HTML wrapper layout is disclosed with each figure. The surrounding reading notes are our own.

Text reading scope & known omissions
  • Abstract; Sections 1–2, Introduction and Related Work
  • Sections 3.1–3.5, benchmark setup, schemas, stratification, scoring, language and simulation
  • Sections 4.1–4.5, scaling, language, task sharing, model comparison and simulation diagnostics
  • Section 5, Conclusion and Limitations; References
  • Appendices A–D, broader impacts, baseline scope, implementation consistency and resources
  • Appendices E–H, per-task model, language, scaling and transfer results
  • Appendices I–N, dataset statistics, observation/action layout, algorithms, scoring rubrics, limitations and OOD audit

Outside the original text pass

  • Text extraction does not reconstruct figure images; inspect the retained PDF for figures and equation/table layout.
  • Separate supplemental material availability has not been fully verified.
  • Version identity: this report uses arXiv:2603.28545v2 (1 July 2026). The official abstract metadata and embedded PDF title retain the catalog title, while PDF page 1 reads "ManipArena: A Controlled Benchmark for Diagnosing Generalization in Real-Robot Manipulation". This discrepancy is preserved; its cause is unknown. The explicit version PDF is byte-identical to the originally retained unversioned download, whose URLs remain in unchanged extracted-text page labels. The v1 PDF was not read and no scientific comparison between versions is claimed.
  • Identity notes: the catalog title is "ManipArena: Comprehensive Real-world Evaluation of Reasoning-Oriented Generalist Robot Manipulation". The supplied official version-history and PDF-link fragments establish the v1-to-v2 identifier chain and author continuity; the supplied acquisition record establishes byte identity with the explicit v2 PDF. Page 1 independently verifies the observed title, v2 date and all 27 catalog authors. The v1 metadata has a shorter byline; this report does not claim to have read its PDF or explain the metadata/title discrepancy.
  • The complete supplied 20-page text, including all appendices and references, was read. Original Figures 1–5 and Tables 1–15 were visually inspected; six final original crops were inspected. The extraction limitation above was addressed by the PDF visual pass.
  • External code, dataset files, full schema assets, separate supplements and official web pages beyond the supplied identity fragments were not inspected. No experiments were reproduced. Dataset licensing and release accessibility were not independently verified.

The visual audit above records the subsequent illustrated pass.

8.4 Traceable evidence

e01PDF p. 1, visible title, complete byline, affiliations and arXiv marginInspect

The visible title is ManipArena: A Controlled Benchmark for Diagnosing Generalization in Real-Robot Manipulation. The margin identifies 2603.28545v2, 1 Jul 2026; the page lists 27 authors and five affiliations.

Go to primary source ↓
e02PDF p. 1, Abstract/Introduction; p. 5, Figure 2, Table 2 and Sections 3.1–3.2Inspect

The benchmark fixes physical conditions, uses tabletop and mobile platforms at 20 Hz, and reports 20 tasks, 10,812 trajectories, 13.5M frames and about 188 robot hours.

Go to primary source ↓
e03PDF p. 6, Section 3.3, Equation (2); p. 19, Algorithm 2, lines 1–10Inspect

The tabletop protocol assigns trials 1–4 to ID, 5–8 to controlled shifts and 9–10 to the strongest held-out condition, then resets, executes and scores the policy.

Go to primary source ↓
e04PDF p. 6, Section 3.4, Equations (3)–(4); p. 19, Table 14Inspect

Subgoal weights total ten points per trial; aggregate score and success at score ≥9 are distinct. Table 14 specifies task rubrics, including fruit penalties.

Go to primary source ↓
e05PDF p. 5, Section 3.2; p. 6, Sections 3.2/3.5; pp. 15–16, Appendices J/K; p. 18, Table 13 and Algorithm 1Inspect

Task configurations precede resets, teleoperation and three-level annotation. Full schema distributions are external to the PDF. LeRobot v2.1 vectors contain 56 tabletop or 62 mobile fields; Table 13 identifies positions, velocities, currents, grippers and mobile extras.

Go to primary source ↓
e06PDF p. 13, Appendices B–C; p. 14, Table 6; p. 18, Appendix M, Instruction interfaceInspect

Within-study controls preserve interfaces, prompts and rubrics. Upstream pretraining differs. OpenPI/LeRobot, Fast-WAM and DreamZero adapter conventions are described, but a complete shared neural objective or optimizer recipe is not given.

Go to primary source ↓
e07PDF p. 7, Figure 3 and Section 4.1; p. 16, Table 10 including dagger footnoteInspect

The five-task 200/task recipe scores 312 with 42% SR; full 60k scores 277/34%, and full 120k scores 248/26%. The tasks are glasses, blocks, wire, button and fruits. All non-dagger recipes train 60k steps.

Go to primary source ↓
e08PDF pp. 7–8, Section 4.2 and Table 4; p. 16, Table 9, category blocks and total rowInspect

Training-language variants all use L1 at evaluation. Table 9 prints L1/L2/L3 totals 702.5/867.25/801.75 and SR 20.0/32.7/35.3; Table 4 prints execution 421/535/432 and semantic 282/333/370.

Go to primary source ↓
e09PDF p. 7, Table 3; p. 8, Section 4.3Inspect

Unified minus specialist total score is +76.3 for π0.5 and −86.0 for WALL-OSS-0.5. The pretraining-provenance interpretation is explicitly correlational.

Go to primary source ↓
e10PDF p. 8, Section 4.4; p. 9, Table 5 and Figure 5Inspect

Table 5 prints aggregate model scores and SRs, including WALL-OSS specialist 951/38%, π0.5 unified 703/20%, and Wall-WM 868/31.33%. Section 4.4 reports 950.8 for the best score and describes the seven-configuration study.

Go to primary source ↓
e11PDF p. 6, Section 3.5; p. 9, Section 4.5; p. 14, Appendix H; p. 17, Table 11Inspect

Three real/sim/mixed checkpoints are evaluated on the real robot. press_button scores/SRs are 31/0%, 65/50%, 40/20%; pick_fruits 78/50%, 51/30%, 68/50%; put_blocks 83/70%, 85/70%, 85/70%.

Go to primary source ↓
e12PDF p. 15, Table 8, Mobile Manipulation block and captionInspect

Only task-specific π0.5 has mobile results: five scores 4.5, 0, 8, 14 and 10, all at zero success; other configurations are marked unavailable.

Go to primary source ↓
e13PDF p. 17, Appendix M, stochasticity, cost/coverage and platform/environment scope; p. 18, Instruction interface; pp. 9, 15–17, Tables 5 and 8–11, uncertainty-reporting scopeInspect

The authors discuss physical stochasticity, limited model/mobile/simulation coverage, controlled environments and possible benefits of alternative prompts. Printed result tables do not report confidence intervals.

Go to primary source ↓
e14PDF p. 13, Appendix D; p. 14, Table 7, π0.5 rows and captionInspect

Approximate fine-tuning resources are 30k steps, 8×A800, 6h/48 GPU-hours for a specialist and 60k, 8×A800, 20h/160 GPU-hours for unified π0.5; upstream pretraining and physical labor are excluded from GPU-hour accounting.

Go to primary source ↓
e15PDF p. 6, Section 3.3; p. 18, Appendix N; p. 20, Table 15Inspect

Seven tasks introduce new objects at T9–T10; eight use configuration OOD. Examples include new sunglasses for put_glasses and held-out button orders for press_button.

Go to primary source ↓
e16PDF p. 8, Table 4 and Section 4.4; p. 9, Table 5; p. 15, Table 8, arrange_cup and model headings; p. 16, Table 9, L1 column and total rowInspect

Tables 5/8 show an extra Wall-WM configuration beyond the prose count. Table 9 L1 arrange_cup is 18 rather than Table 8 unified 47; printed L1 task scores do not sum to the stated 702.5 total. Table 4 L1 execution SR is 15, versus 16 in Table 5.

Go to primary source ↓
e17PDF p. 16, Appendix K, schema-release statement; p. 14, Tables 6–7; p. 19, Table 14, mobile rowsInspect

The PDF points outside itself for full schema distributions, provides only partial training configuration details and summarizes several mobile rubrics without every numerical stage weight.

Go to primary source ↓

8.5 Primary sources

Scroll across the image to inspect details. Press Esc to close.