PAPER REPORTENAll readings ↗

FolDeX: A Physical-World Benchmark for Long-Horizon Robotic Manipulation of Deformable Objects

English reading report: Method, equations, original figures, experiments and reproducibility.

Authors: Chenhuan Liu; Yi Xu; Feng Wu; Hanyang Wang; Wenxiao Kuai; Weihao Ding; Shan Wang; Yang Liu; Shuyong Gao; Wenqiang Zhang

Affiliations: Fudan University; AI Research Center, Midea Group (Shanghai) Co., Ltd.; Carnegie Mellon University

Source: 2609.10243 ↗ · Catalog record

Reading: 13 / 558 · 5 original figures & tables · ~17 min ·

1. Paper overview

In one sentence: FolDeX turns complete physical garment folding into a controlled test of robot-data reuse, with promising recovery-augmentation results but incomplete scoring and transfer specifications. e-interfacee-trackse-baselinese-recoverye-missing-appendixe-transfer

At a glanceWhat to know
Research problem
Source description

A policy can succeed in simulation or on isolated grasping and folding stages yet fail when retrieval, flattening, folding, and stacking must all work consecutively. Deformation, self-occlusion, contact changes, and accumulated errors make garments a demanding test. The authors also argue that different resets, data mixtures, latency handling, and baseline implementations can distort physical-robot rankings. FolDeX therefore targets complete closed-loop episodes under a shared evaluation protocol. e-motivatione-platforme-interface

Core mechanism
Source description

A real-robot dataset and benchmark with 2,000+ hours, 20+ tasks, and 10+ embodiments, centered on garment folding. Its common schema records synchronized vision, proprioception, timestamped actions, language instructions, and episode metadata. e-data

A key reported resultRecovery-augmented complete garment folding: 95.00%; 82.53; 1:58 (minutes:seconds). Category success: 90.00%, 90.00%, 100.00%, 100.00%, respectively.

Average success rate; FoldScore; reported average time. One RTC π0 policy trained on demonstrations plus recovery data; Shirt, Skirt, Pants, and Towel under the complete physical protocol; default 30 trials per category.

Demo-only RTC multi-task: 80.75% and 75.59. Calculated differences: +14.25 percentage points and +6.94 FoldScore points. Measured physical results favor recovery augmentation on average, especially Pants and Towel. The comparison adds data and does not isolate recovery content from extra training experience. e-protocole-baselinese-recovery

Reading caution
Reader analysis

Cross-task and cross-embodiment studies report interference and forgetting under joint or continued training; moderate scene changes appear less disruptive. These are preliminary qualitative observations without axis-specific numerical tables, data budgets, or uncertainty, not quantified transfer laws. e-transfer

Core contributions

  • Source description

    A real-robot dataset and benchmark with 2,000+ hours, 20+ tasks, and 10+ embodiments, centered on garment folding. Its common schema records synchronized vision, proprioception, timestamped actions, language instructions, and episode metadata. e-data

  • Author claim

    Four data-reuse tracks separate recovery augmentation, cross-task transfer, scene changes, and embodiment changes. FoldChallenge is described as an operational external-policy evaluation platform with held-out physical objects, controlled initializations, and auditable rollouts. e-platforme-tracks

  • Reader analysis

    Demonstration-only π0 reference comparisons, a recovery-augmented comparison, and FoldScore connect execution success to final garment quality and efficiency. These are benchmark reference results, not a new predictive world-model architecture. e-metrice-baselinese-recovery

Figure 1. The benchmark connects reusable robot experience to standardized physical evaluation. Original paper, p. 2 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Read downward through three different roles. The upper cartoons motivate the sim-to-real and rigid-to-deformable gaps; they are not measured comparisons. The middle row shows what changes in each research track: recovery examples, task content, scene appearance or layout, and robot embodiment. Along the bottom, solid arrows run from benchmark data to the evaluation pipeline, rollout repository, and leaderboard. A dashed arrow returns from the repository toward the data side, illustrating reuse of deployment experience. The caption and track descriptions support this reading. These are data and evaluation relationships, not neural-network layers or a future-video-to-action computation graph. e-platforme-trackse-baselinese-visual-disclosure

What it supports. FolDeX's proposed contribution is control over what data are reused and how the resulting policies are tested. The repository and leaderboard complete an evaluation cycle around physical execution. The authors describe held-out objects and controlled initialization as safeguards for comparing submitted policies under consistent conditions.

Where the evidence stops. This is a conceptual overview, not proof of platform availability or a measured simulation gap. Its leaderboard uses rounded values; use Table 2 for precise results. The authors disclose AI-assisted visual drafting, and the PDF retains contradictory public-access/anonymity language.

2. Motivation

2.1 The problem and the proposed response

Source description

A policy can succeed in simulation or on isolated grasping and folding stages yet fail when retrieval, flattening, folding, and stacking must all work consecutively. Deformation, self-occlusion, contact changes, and accumulated errors make garments a demanding test. The authors also argue that different resets, data mixtures, latency handling, and baseline implementations can distort physical-robot rankings. FolDeX therefore targets complete closed-loop episodes under a shared evaluation protocol. e-motivatione-platforme-interface

2.2 What this reading follows

Folding a garment is a chain of dependent physical decisions: a poor retrieval can make flattening harder, and a misplaced fold can spoil the final stack. FolDeX evaluates that complete chain and organizes real-robot experience around recovery, task, scene, and embodiment changes. Its π0 comparisons show why both the training data and the metric matter: recovery augmentation improves average success, while task-specific and multi-task policies trade places when ranked by success versus FoldScore. Read the figures as a benchmark design and the tables as preliminary physical evidence. The supplied version leaves the detailed quality rubric, some execution settings, and quantitative transfer comparisons unresolved. e-interfacee-trackse-baselinese-recoverye-missing-appendixe-transfer

3. Research context

We place the paper in the collection through its world–action interface. The catalog labels and the reading’s assessment are shown separately.

Catalog dimensionRecorded classification
Major categoryNot assigned
ArchitectureNot assigned
Prediction paradigmNot assigned
QuadrantNot assigned

This table preserves the labels recorded at reading time. The current major category is Benchmarks & simulators. View the current classification.

3.1 Evidence-based assessment

Classification assessment not applicable

Reader analysis

The recorded fields remain Not assigned. FolDeX contributes a dataset, physical benchmark, metric, and reference-policy comparisons. A One Model/Two Model or joint-future-action/inverse-dynamics quadrant does not apply to this benchmark itself. A single multi-task π0 baseline is not evidence for an integrated world-action architecture, and mentioning WAMs in the motivation does not establish one. e-motivatione-interfacee-rtce-baselines

This is the collection’s architectural analysis, not a new related-work survey. Benchmark comparisons and their protocols appear in Section 6.

4. Problem formulation

4.1 Inputs and outputs

InputsOutputs
  • Synchronized RGB from two wrist cameras and one external workspace camera
  • Natural-language task instruction and robot proprioceptive state
  • Policy prediction: a future trajectory of robot end-effector poses
  • Evaluation: complete-task success, completion time, final-state neatness, and FoldScore

4.2 Equations and their role

FoldScore=35Qˉ+35R+30max(0,1tTmax)\mathrm{FoldScore}=35\bar{Q}+35R+30\max\left(0,1-\frac{t}{T_{\max}}\right)
Equation (1): Q̄ is average normalized final-state quality in [0,1]; R is complete-task success rate in [0,1]; t is mean completion time over successful trials; Tmax is maximum episode duration. The time contribution is zero if no trial succeeds. Higher scores reward quality, success, and speed. The source supplies neither the promised neatness normalization mapping nor a numerical Tmax in this PDF. e-metrice-missing-appendix

5. Method in detail

5.1 Follow the episode before interpreting the model

Reader analysis

Begin with the state the robot actually receives: the garment is in a basket, not already arranged for folding. Retrieval creates a new table configuration, which determines what flattening can accomplish; flattening then constrains the folds and the final stack. The policy repeatedly receives camera observations and proprioception and predicts future end-effector poses. This source-defined interface makes physical feedback central, but it does not introduce a predictive world model. The tutorial implication is that a local manipulation success cannot substitute for completion of the whole chain. Figure 2 explains why the trial definition matters: the photographs sample stages and task diversity, whereas the success rule on page 6 requires an entire episode without human intervention or unrecoverable failure. e-interfacee-protocol

Figure 2. Complete folding spans multiple stages within a broader collection of tasks and robots. Original paper, p. 4 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Use the source caption to interpret the four rows. The top row samples stages of a long folding episode; the second shows garment diversity; the third adds other deformable and rigid-object tasks; the bottom varies robot embodiment and scene. The adjacent task definition supplies the full sequence: retrieve from a basket, flatten, fold, and place or stack. The montage should therefore be read together with that definition, rather than as four independent benchmark successes. At inference, the policy uses two wrist-camera views, an external workspace view, an instruction, and proprioception to predict future end-effector poses. e-interfacee-datae-protocole-baselinese-transfer

What it supports. The benchmark tests whether progress survives a sequence of contacts and changing garment configurations. A successful isolated fold is insufficient for the complete task. The surrounding task and embodiment diversity motivates transfer research, while the main numerical tables focus on four garment categories.

Where the evidence stops. Still images show examples of states and setups, not uninterrupted autonomous rollouts. They do not establish success rates or demonstrate cross-embodiment transfer. The paper describes the camera/action interface but omits detailed control frequency and trajectory-horizon settings.

5.2 Separate delayed execution from recovery experience

Reader analysis

Two changes address different problems. Training-time RTC simulates inference delay and conditions the policy on action prefixes already committed for execution. Its reference comparison changes the training configuration while keeping the multi-task, demonstration-only setting. Recovery augmentation instead changes the experience available for learning: operators collect corrections from deployed-policy states, and the track combines those trajectories with independent demonstrations for a separately initialized policy. Table 3 applies that idea to a single RTC π0 trained across four garments. A plausible reader interpretation is that recovery examples cover states missing from clean demonstrations, but the paper does not measure that coverage or isolate it from additional data volume. The experiments support a useful combination, while a budget-matched control is still needed to identify why it helps. e-rtce-trackse-baselinese-recovery

Figure 3. The dataset is heterogeneous and visibly uneven across tasks and embodiments. Original paper, p. 6 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Both horizontal axes count trajectories, not hours or success percentages. The left panel groups records by embodiment, with orange deformable and blue rigid portions stacked where both appear. The right panel groups trajectories by task using the same colors. Aloha dominates the embodiment panel, while Fold shirt is the largest task category. Keep the panels separate: they are marginal summaries, not a task-by-robot matrix. The common data schema described on page 4 includes object-instance, scene, embodiment, and completion metadata, which would be needed to construct controlled subsets from these aggregate distributions. e-datae-baselinese-recoverye-transfer

What it supports. The collection spans several task families and robots, but its diversity is not balanced coverage. That distinction matters for interpreting data reuse: a gain from mixing datasets could reflect which categories dominate the mixture. This is a reader inference motivating controlled sampling, not a measured explanation of the reported transfer failures.

Where the evidence stops. The bars lack exact value labels and do not specify the actual training mixtures for Tables 2–3. Do not convert approximate bar lengths into exact split sizes or infer per-embodiment performance from these dataset counts.

5.3 Read FoldScore as a tradeoff with missing calibration

Reader analysis

Equation (1) assigns 35 times average normalized quality, 35 times success rate, and up to 30 points for successful-trial speed. A policy can therefore lose in binary success yet win the composite, as RTC Multi does against RTC Single in Table 2. This is not enough to identify the cause of that ranking: the table does not expose the underlying mean quality and time components. The paper also refers the six-level neatness rubric and its normalization to an appendix absent from the supplied PDF. Consequently, do not assume that a neatness level is simply divided by five. Read the printed FoldScores as reported outcomes, preserve the separate success rates, and require the missing mapping and episode timeout before independently reconstructing or recalibrating the metric. e-metrice-baselinese-missing-appendix

5.4 Training and inference

During training

Source description

The main baselines use π0 with expert demonstrations: standard multi-task training, RTC multi-task training, and RTC task-specific training. Training-time real-time chunking simulates inference delay and conditions on already committed action prefixes, avoiding inference-time inpainting overhead. Recovery training adds trajectories from all four evaluated garment categories. e-rtce-baselinese-recovery

Reader analysis

All models are reported as trained on 16 GPUs with batch size 16. GPU models, training duration, optimizer settings, explicit loss equations, and frozen-module choices are not specified here; the paper should not be treated as a complete π0 implementation recipe. e-protocole-rtc

During inference

Reader analysis

At each inference step, current camera observations, instruction, and proprioception produce future end-effector poses for physical execution. The benchmark emphasizes closed-loop feedback. It does not specify a new latent state, future-video generator, inverse-dynamics module, or test-time imagination loop; nor does it give the trajectory horizon or control frequency. e-interfacee-rtc

5.5 Implementation flow

  1. Define a dependent physical episode

    Start with a garment in a basket. Retrieve it onto the table in a random configuration, flip and flatten it, perform multiple folds, then place or stack it. Earlier physical errors propagate into later stages; stage completion alone is insufficient. e-interface

  2. Organize heterogeneous experience

    Record task, object category and instance, scene, embodiment, operator, data source, completion status, and failure or termination reason. Platform diversity includes Aloha, YAM, Piper, Astribot, Franka, UR5, a mobile dual-arm platform, and AgiBot G2; a shared record format does not itself align their action spaces. e-datae-transfer

  3. Construct the recovery track

    Deploy a shared policy, collect human intervention and recovery trajectories, and combine them with independently collected demonstrations to train a separately initialized policy. The evaluation tests subsequent autonomous behavior, not success achieved through intervention during the scored episode. e-trackse-protocol

  4. Hold evaluation conditions fixed

    Use the same tasks, garment instances, initialization procedure, maximum episode duration, and robot execution protocol. Unless otherwise specified, evaluate each garment category in 30 independent physical trials with randomized configurations. Success requires the whole sequence without human intervention or unrecoverable failure. e-metrice-protocol

  5. Assess outcome quality

    Neatness has six levels, 0–5, based on final shape, compactness, flatness, structural stability, and edge alignment. At least two evaluators independently assess qualitative scores; the same evaluators assess all methods and a score is recorded only on agreement. e-metrice-protocol

6. Experiments & results

FolDeX evaluates complete physical garment-folding episodes and asks whether expensive robot experience can be reused across recovery states, tasks, scenes, and embodiments. Its reference policies use π0. Recovery-augmented training raises reported average success from 80.75% to 95.00%, but transfer studies remain preliminary and the missing quality-scoring appendix prevents independent reconstruction of FoldScore from the supplied paper alone.

Source and visual limitations
Reader analysis

This benchmark paper contains no new neural-network architecture diagram or dedicated architecture-ablation figure. Figure 1 supplies the protocol overview, while Table 2 supplies the RTC/task-sharing diagnostic comparison and Table 3 the recovery-data comparison. Cross-task, cross-scene and cross-embodiment findings are qualitative without dedicated quantitative plots or tables. The promised appendix with the quality rubric is absent, so no rubric visual can be supplied. e-platforme-rtce-baselinese-recoverye-transfere-missing-appendix

6.1 Read the original evidence

Table 3. Recovery-augmented training improves the reported average over the demonstration-only reference. Original paper, p. 7 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. This table evaluates one multi-task π0 policy trained with RTC on demonstrations plus recovery trajectories from all four garments. Read each row across success, average time in minutes:seconds, and FoldScore. The metric definition uses time over successful trials, so time should be interpreted alongside success rather than as a failure-inclusive duration. To assess augmentation, compare this table with the RTC Multi column of Table 2, not with task-specific training or the supplementary pre-flattened Shirt experiment. The recovery track collects human corrections during earlier deployment, then evaluates the resulting policy under the complete-task rule without human intervention. e-recoverye-baselinese-trackse-protocole-metrice-missing-appendixe-supplementary

What it supports. Average success rises from 80.75% to 95.00%, and FoldScore from 75.59 to 82.53. Pants and Towel reach 100.00% success, while Shirt and Skirt remain at 90.00%. The reported aggregate improvement is strong, but it is not a uniform gain: Shirt success falls from Table 2's printed 93.00%.

Where the evidence stops. Adding recovery trajectories changes both data content and data quantity. Without an equal-budget extra-demonstration control, the result cannot isolate recovery-specific benefit. Finite trial counts, missing uncertainty, and unavailable scoring details further limit claims of robust superiority.

6.2 Results and evaluation conditions

Task & protocolReported resultComparison & interpretation
Recovery-augmented complete garment folding

One RTC π0 policy trained on demonstrations plus recovery data; Shirt, Skirt, Pants, and Towel under the complete physical protocol; default 30 trials per category.

95.00%; 82.53; 1:58 (minutes:seconds). Category success: 90.00%, 90.00%, 100.00%, 100.00%, respectively.

Average success rate; FoldScore; reported average time

Demo-only RTC multi-task: 80.75% and 75.59. Calculated differences: +14.25 percentage points and +6.94 FoldScore points.

Measured physical results favor recovery augmentation on average, especially Pants and Towel. The comparison adds data and does not isolate recovery content from extra training experience. e-protocole-baselinese-recovery

Training-time RTC and task sharing

Table 2; expert demonstrations only; four garment categories under the complete physical protocol.

RTC Multi: 80.75 / 75.59; Standard Multi: 72.50 / 72.30; RTC Single: 82.50 / 71.69.

Average success rate (%) / FoldScore

RTC improves multi-task success by a calculated 8.25 percentage points. Single-task RTC leads average success; multi-task RTC leads FoldScore.

Metric choice changes the ranking. Sharing is task-dependent: Towel success is 76.67% for RTC Multi versus 50.00% for RTC Single, while Pants is 63.33% versus 80.00%. e-protocole-rtce-baselines

Supplementary Shirt folding from a pre-flattened state

Reproduced π0.5 baseline; folding after pre-flattening, not the full retrieval-to-placement episode.

100.0%; 41 s; generally the lowest non-zero quality tier.

Success rate; mean completion time; qualitative final-state quality

The supplementary inference-time RTC π0 result is 42.9% on the complete Shirt task, a different protocol.

These reported results cannot establish a direct model ranking because their initial states and episode scope differ. e-supplementary

6.3 Ablations and diagnostic examples

Read component removals and qualitative examples within their stated evaluation conditions.

Table 2. Training-time chunking and task sharing affect success and composite quality differently. Original paper, p. 6 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Each garment has a success-rate row and a FoldScore row. All columns use π0 and expert demonstrations; RTC means real-time chunking during training. Compare RTC Multi against Standard Multi to examine the RTC configuration within joint garment training. Compare RTC Multi against RTC Single to examine shared versus category-specific policies. The bottom two rows summarize the four categories, but the task rows expose tradeoffs hidden by those averages. For example, RTC Multi improves Pants success relative to Standard Multi while its Pants FoldScore is lower. FoldScore additionally incorporates final-state quality and successful-trial completion time, as defined on page 5. e-rtce-baselinese-protocole-metrice-missing-appendix

What it supports. RTC Multi reaches 80.75% average success versus 72.50% for Standard Multi, a calculated 8.25-percentage-point increase. RTC Single achieves the highest average success, 82.50%, but RTC Multi has the highest FoldScore, 75.59. Thus neither training choice nor metric produces a universally dominant policy across garments.

Where the evidence stops. This is a configuration comparison, not a fully controlled architecture ablation. No uncertainty is reported. Shirt's printed 93.00% needs raw-count clarification under the stated default of 30 trials; the absent quality rubric also limits independent FoldScore auditing.

7. Analysis & limitations

7.1 What the evidence leaves open

Reader analysis

Cross-task and cross-embodiment studies report interference and forgetting under joint or continued training; moderate scene changes appear less disruptive. These are preliminary qualitative observations without axis-specific numerical tables, data budgets, or uncertainty, not quantified transfer laws. e-transfer

Source description

The authors acknowledge costly physical evaluation, limited model coverage, and non-exhaustive track studies. Current observations lack explicit tactile sensing, which they propose adding later. e-limits

Reader analysis

Confidence intervals and raw trial outcomes are absent. Table 2 prints 93.00% for RTC Multi Shirt despite the default 30-trial protocol; the exact success count or aggregation is unresolved. Reported values are retained without silently changing 93.00 to 93.33. e-protocole-baselines

7.2 Questions for discussion

  1. Would recovery augmentation retain its advantage over an equal budget of additional demonstrations?
  2. How stable are FoldScore rankings across garment instances, raters, and episode timeouts?
  3. Can embodiment transfer preserve complete-episode performance once action conventions are explicitly aligned?

8. Reproducibility audit

8.1 Requirements and known gaps

Reader analysis

Reconstruction needs the exact train/test object-instance partitions, demonstration/recovery mixtures, collection procedure, and embodiment-specific action conventions. The source describes these categories and controlled evaluation intentions but supplies no complete split manifest, dataset license, or versioned release specification. e-platforme-datae-tracks

Reader analysis

Obtain the missing neatness rubric, normalization mapping, numerical episode timeout, and per-trial component scores before recomputing FoldScore. Agreement-only qualitative scoring also needs a documented disagreement-resolution procedure. Success rates alone cannot recover missing quality and timing components. e-metrice-protocole-missing-appendix

Reader analysis

Proposed checks: compare recovery data against an equal additional demonstration budget with fixed optimization and physical evaluation; separately test training-time RTC at matched measured latency while logging success, quality, and successful-trial duration. Both checks require raw counts and repeated training runs. e-rtce-baselinese-recoverye-metric

8.2 Proposed reproduction checks

The following checks are proposals motivated by the paper. They have not been run as part of this reading.

Check 1: Does recovery content beat an equal amount of new demonstration data?

Reader-proposed, not performed: use the same π0 initialization, RTC settings, four-garment base demonstrations, optimization steps, and held-out object protocol. Compare the base set, base plus recovery trajectories, and base plus newly collected demonstrations matched to the added recovery data in both sampled transitions and collection budget where feasible; report any mismatch. Train repeated seeds and evaluate at least the source's 30 randomized trials per category. Log exact success counts, stage failures, and raw quality/time components. A reproducible recovery advantage over the matched extra-demonstration condition, especially on Pants and Towel, would support a recovery-specific explanation; gains explained equally by either data addition would weaken it. e-trackse-protocole-baselinese-recoverye-metric

Check 2: Does RTC's advantage persist under measured latency and auditable scoring?

Reader-proposed, not performed: repeat Standard Multi versus training-time RTC Multi with identical demonstration data, training budget, physical robot, garment instances, resets, and episode timeout. Measure deployed inference delay and evaluate both policies under the same delay schedule, including controlled lower and higher delays. Record committed action prefixes, raw success counts, successful-trial times, and independently rated final states. Obtain the missing original rubric and normalization before reconstructing FoldScore; until then report components only. An RTC advantage that grows with delay would support the timing rationale. If the gain disappears at matched latency or depends on unexplained scoring differences, the original configuration comparison would not isolate that mechanism. e-rtce-baselinese-protocole-metrice-missing-appendix

8.3 Reading coverage

Visual audit: All nine original PDF pages were rendered and visually inspected, including the title/author/version page; the benchmark and task figures; the metric equation and Table 1; the trajectory chart and Table 2; Table 3, transfer observations and limitations; and the final disclosure/references pages. All five final original crops were inspected separately. Figure 1's arrow flow was cross-checked against its caption and the track descriptions; it represents benchmark data/evaluation flow, not model inference. Table headers, chart axes and legends are retained, with table-caption definitions explained in the local guides. No appendix is present in this PDF. External videos, platform operation, code, and separate supplements remain uninspected.

PDF pages inspected for this edition: 1, 2, 3, 4, 5, 6, 7, 8, 9. Appendix coverage: not present.

Original figures and tables remain the work of the source’s authors. Extractions preserve their scientific content; any HTML wrapper layout is disclosed with each figure. The surrounding reading notes are our own.

Text reading scope & known omissions
  • PDF p. 1: title, authors, affiliations, version, Abstract
  • PDF pp. 1–3: Introduction and contributions
  • PDF pp. 3–4: Related Work
  • PDF p. 4: FolDeX; Design Principles; Task Suite and Policy Interface; Data Collection and Organization
  • PDF pp. 4–5: Benchmark Tracks
  • PDF p. 5: Evaluation Metrics and Equation (1)
  • PDF pp. 6–7: Experiments; Baselines; Supplementary baselines; Recovery-Data Utilization; Preliminary Observations on Other Transfer Axes
  • PDF p. 7: Discussion and Limitations; Conclusion
  • PDF pp. 8–9: Use of Generative AI; References

Outside the original text pass

  • Text extraction does not reconstruct figure images; inspect the retained PDF for figures and equation/table layout.
  • The complete supplied nine-page PDF was read. Its title and ten authors match the catalog; the observed version is arXiv:2609.10243v1 [cs.RO], 9 September 2026. The title page also prints a 2027 AAAI copyright notice; this does not establish a later edition or venue acceptance. No other revision was supplied or compared.
  • The text extraction did not reconstruct figure images; original PDF pages, Figures 1–3 and Tables 1–3 were separately inspected.
  • Separate supplemental material availability has not been fully verified.
  • The scoring rubric and neatness-to-quality mapping promised in the Appendix on p. 5 are absent from this PDF, which ends with references on p. 9.
  • Code, external platform operation, downloadable data, and rollout videos were not inspected; no experiments were reproduced. The PDF claims public access but retains review-anonymity language saying identifying information and links are omitted.

The visual audit above records the subsequent illustrated pass.

8.4 Traceable evidence

e-identityPDF p. 1, title/author block, affiliation lines, arXiv margin stamp and copyright footerInspect

The exact FolDeX title and all ten catalog authors are printed. Affiliations are Fudan University; AI Research Center, Midea Group (Shanghai) Co., Ltd.; Carnegie Mellon University. The stamp identifies arXiv:2609.10243v1, 9 Sep 2026; the footer prints copyright 2027 AAAI.

Go to primary source ↓
e-motivationPDF p. 1, Abstract and Introduction; PDF p. 2, Introduction continuationInspect

The authors motivate full physical evaluation through sim-to-real and rigid-to-deformable gaps, dependent bimanual stages, and sensitivity to baseline engineering and evaluation settings.

Go to primary source ↓
e-platformPDF p. 2, Figure 1 and caption, Introduction right column; PDF p. 3, opening lines; PDF p. 1, AbstractInspect

Figure 1 organizes four data-reuse axes and a data-to-evaluation-to-rollout-to-leaderboard workflow. The text claims an operational public platform with held-out objects and controlled physical execution. Review-anonymity language on pp. 2–3 says identifying information and access links are omitted, although p. 1 prints a platform URL.

Go to primary source ↓
e-interfacePDF p. 4, Figure 2 and caption; Design Principles; Task Suite and Policy InterfaceInspect

The interface supplies synchronized RGB from two wrist cameras and one workspace camera, language, and proprioception, and predicts end-effector pose trajectories. Episodes run from basket retrieval through flattening, multiple folds, and final placement/stacking. Figure 2 shows stages, garments, additional tasks, robots, and scenes.

Go to primary source ↓
e-dataPDF p. 4, Data Collection and Organization; PDF p. 6, Figure 3 and captionInspect

The reported dataset spans 2,000+ hours, 20+ tasks and 10+ embodiments. Named platforms and the common modality/metadata schema are described. Figure 3 plots trajectory counts by embodiment and task, with orange deformable and blue rigid categories; Aloha and Fold shirt have the largest bars in their panels.

Go to primary source ↓
e-tracksPDF p. 5, Recovery-data utilization; Cross-task transfer; Cross-scene transfer and generalization; Cross-embodiment transferInspect

Recovery data are collected during a shared policy's deployment and combined with independent demonstrations for a separately initialized policy. Other tracks study garment/rigid transfer, scene adaptation and zero-additional-data generalization, and sharing experience across incompatible embodiments.

Go to primary source ↓
e-metricPDF p. 5, Evaluation Metrics, Final-state neatness, Equation (1) and following definitionsInspect

Neatness levels 0–5 consider shape, compactness, flatness, stability and edge alignment. FoldScore weights normalized quality and success by 35 each and clipped successful-trial time efficiency by 30; no-success time contribution is zero. The rubric and quality mapping are referred to an Appendix. Evaluation factors are held constant.

Go to primary source ↓
e-protocolPDF p. 6, Experiments opening paragraphInspect

Unless specified otherwise, each garment category uses 30 independent randomized physical trials. Success requires the complete sequence without human intervention or unrecoverable failure. Training uses 16 GPUs and batch size 16. At least two consistent evaluators assess qualitative scores independently and record them on agreement.

Go to primary source ↓
e-rtcPDF p. 6, Baselines text, including continuation beneath Table 2Inspect

Reference policies use π0 with expert demonstrations. Multi-task standard and training-time RTC are compared with RTC task-specific policies. Training-time RTC simulates inference delay and conditions on committed action prefixes to avoid inference-time inpainting overhead.

Go to primary source ↓
e-baselinesPDF p. 6, Table 2, all task rows and Average SR/FoldScore rows; PDF p. 7, Baselines continuationInspect

RTC Multi/Standard Multi/RTC Single averages are 80.75/72.50/82.50% success and 75.59/72.30/71.69 FoldScore. RTC Multi Shirt is printed as 93.00%. Pants success is 63.33/50.00/80.00%; Towel is 76.67/50.00/50.00%. Pants FoldScore is 70.52/73.33/76.23, showing that RTC need not improve both metrics for each task.

Go to primary source ↓
e-recoveryPDF p. 7, Table 3, all rows and columns, caption; Recovery-Data UtilizationInspect

A single RTC π0 trained with demonstrations and recovery data obtains Shirt/Skirt/Pants/Towel success of 90/90/100/100%, times 2:29/1:54/1:48/1:40, and FoldScores 76.88/81.25/85.50/86.50. Reported averages are 95.00%, 1:58 and 82.53; the text compares against demo-only 80.75% and 75.59.

Go to primary source ↓
e-supplementaryPDF p. 7, Supplementary baselinesInspect

π0.5 is tested only after pre-flattening and achieves 100.0% success in 41 s, generally at the lowest non-zero quality tier. π0 with inference-time RTC scores 42.9% on the full Shirt task. The authors explicitly exclude these incompatible protocols from a direct main comparison.

Go to primary source ↓
e-transferPDF p. 7, Preliminary Observations on Other Transfer Axes, both columnsInspect

The authors qualitatively report action interference under heterogeneous joint training, forgetting under continued training, more severe embodiment problems, and greater robustness to lighting/background/moderate layout changes. They label these observations preliminary and give no transfer-axis numerical table.

Go to primary source ↓
e-limitsPDF p. 7, Discussion and LimitationsInspect

Time and hardware costs restrict evaluated models. Track studies are non-exhaustive. The current benchmark uses visual observations and proprioception without explicit tactile sensing; tactile feedback is future work.

Go to primary source ↓
e-missing-appendixPDF p. 5, two Appendix references in Evaluation Metrics; PDF pp. 8–9, document endingInspect

The scoring rubric and neatness normalization are promised in an Appendix, but the supplied PDF ends with the generative-AI disclosure and references and contains no appendix or numerical episode timeout.

Go to primary source ↓
e-benchmark-comparisonPDF p. 5, Table 1 and captionInspect

The source compares benchmark scope, physical versus simulation evaluation, complete long-horizon episodes, community evaluation, and reported signals. This is a descriptive comparison of benchmark designs, not a common-protocol performance ranking.

Go to primary source ↓
e-visual-disclosurePDF p. 8, Use of Generative AIInspect

The authors disclose generative-AI use for language editing and visual drafting and state that they verified the technical content, results, and final figures.

Go to primary source ↓

8.5 Primary sources

Scroll across the image to inspect details. Press Esc to close.