PAPER REPORTENAll readings ↗

CoWAM: Coordination Contracts for Selective Policy Intervention with WAMs

English reading report: Method, equations, original figures, experiments and reproducibility.

Authors: Shuaijun Liu; Qifu Wen; Shuyang Hao; Qi Luo; Chenglong Zhang; Feiyang You; Chengyu Wu; Ningxin Su

Affiliations: The Hong Kong University of Science and Technology (Guangzhou); Boston University; Shanghai Jiao Tong University

Source: 2608.02578 ↗ · Catalog record

Reading: 96 / 558 · 6 original figures & tables · ~19 min ·

1. Paper overview

In one sentence: CoWAM converts predicted bimanual futures into conservative action overrides through coordination contracts, trading some available rescues for fewer unsupported interventions. e-probleme-interfacee-contracte-gatese-evente-naturale-ablation

At a glanceWhat to know
Research problem
Source description

A plausible predicted rollout can conceal mistimed contacts, incompatible arm roles or converging collisions. CoWAM asks when predicted consequences justify replacing Policy Top-1, balancing coordination repairs against damage to an already valid nominal action. e-probleme-gates

Core mechanism
Source description

Coordination contracts combine noncompensable typed predicates, event-conditioned learned evidence and calibrated intervention gates. e-contracte-verifiere-gates

A key reported resultNatural closed-loop success across eight RoboTwin 2.0 tasks: 151/240 successes (62.9%); 95.4% coordination validity; 3.3% false and 0.4% harmful intervention.

Task success; coordination validity; false/harmful intervention rates. 30 paired held-out seeds per task; 240 simulated episodes per method; strict simulator completion within a fixed horizon.

Selective Control: 128/240 (53.3%); Policy Top-1: 96/240 (40.0%); offline oracle: 176/240 (73.3%). The reported gain over Selective Control is 9.6 percentage points, with 32 versus 9 discordant pairs and reported exact paired p=4.3×10⁻⁴. Each task improves by 1–4 episodes; physical-robot transfer is untested. e-naturale-protocol

Reading caution
Reader analysis

Evidence is from RoboTwin 2.0 and held-out seeds of the evaluated task definitions. New-task and physical-robot generalization remain open. The oracle–CoWAM success gap also shows residual selection headroom within the fixed proposal pool. e-protocole-natural

Core contributions

  • Source description

    Coordination contracts combine noncompensable typed predicates, event-conditioned learned evidence and calibrated intervention gates. e-contracte-verifiere-gates

  • Source description

    An outcome-blind same-pool protocol separates selection quality from the availability of successful proposals, with simulator labels obtained only after selectors commit. e-audit

Figure 2. Follow the evidence required to turn a proposed action into an authorized override. Original paper, p. 2 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Read from A to D. Panel A fixes the proposer order: candidate zero is nominal, and other candidates carry paired actions and predicted futures. Panel B activates obligations from task phase and checks every candidate against them; a failure blocks admissibility. Panel C adds event-conditioned ensemble evidence and six conjunctive gates. A high score alone does not authorize execution. Panel D chooses an eligible alternative, preserves a valid nominal action when none qualifies, or invokes the fallback. Finally, follow the separate audit box: decisions are committed before the shared oracle supplies outcome labels. That lower branch is evaluation infrastructure, not an online source of future outcomes. e-interfacee-contracte-verifiere-gatese-audit

What it supports. The mechanism separates proposal availability, contract admissibility and permission to intervene. This explains how a selector can reject an attractive alternative while preserving the nominal action. Because the action–future pool is unchanged, gains are attributed to selection within that pool, subject to the stated information boundary.

Where the evidence stops. The original diagram contains irregular labels and expressions; formal definitions come from Eqs. (2–6). It does not specify the full verifier architecture or fallback implementation, and the gate schematic supplies no safety guarantee.

2. Motivation

2.1 The problem and the proposed response

Source description

A plausible predicted rollout can conceal mistimed contacts, incompatible arm roles or converging collisions. CoWAM asks when predicted consequences justify replacing Policy Top-1, balancing coordination repairs against damage to an already valid nominal action. e-probleme-gates

2.2 What this reading follows

A robot can imagine a plausible future and still choose a poorly coordinated action. CoWAM starts from a frozen WAM's ordered action–future candidates and asks whether there is enough evidence to replace the nominal choice. Its contracts make synchronization, arm roles and collision convergence explicit; a learned verifier then estimates whether an alternative satisfies those obligations without sacrificing progress. The central experiment holds proposals fixed, separating better selection from better generation. This reading follows that decision path, distinguishes event-level validity from complete simulated task success, and examines why removing conservative gates can improve one metric while worsening intervention errors. e-probleme-interfacee-contracte-gatese-evente-naturale-ablation

3. Research context

We place the paper in the collection through its world–action interface. The catalog labels and the reading’s assessment are shown separately.

Catalog dimensionRecorded classification
Major categoryWAMs
ArchitectureDual-system
Prediction paradigmOther mechanisms
QuadrantOutside quadrants

3.1 Evidence-based assessment

Supports the recorded classification

Reader analysis

The frozen WAM proposer and separate verifier/controller support Dual-system, Other mechanisms and Outside quadrants. CoWAM selects existing action–future pairs; it does not introduce a unified future/action generator or inverse-dynamics action decoder. The post-training tag fits an added learned selector, but the source does not establish policy-weight post-training or reinforcement learning. Generalization evidence is limited to the evaluated regimes. e-interfacee-verifiere-traininge-robustness

This is the collection’s architectural analysis, not a new related-work survey. Benchmark comparisons and their protocols appear in Section 6.

4. Problem formulation

4.1 Inputs and outputs

InputsOutputs
  • Current multi-view observation and task-phase/event state
  • An ordered pool of synchronized left-right action chunks, four-step predicted RGB-D futures and predicted proprioception
  • Preserve the nominal action, override with an eligible existing candidate, or invoke the task-defined abstention fallback
  • A committed record of candidate identities, contract outcomes, scores, bounds and decision mode

4.2 Equations and their role

gi=eE(1mt,e+mt,eci,e)g_i=\prod_{e\in\mathcal E}(1-m_{t,e}+m_{t,e}c_{i,e})
Equation (2): the event vocabulary is \(\mathcal E\); \(m_{t,e}\) activates event \(e\), and \(c_{i,e}\) is candidate \(i\)'s binary predicate. Thus \(g_i\) requires every active predicate. e-contract
Ri+=rˉi+κrsir,Ui=uˉiκusiu,Si=Ui+λoo^iλrRi+R_i^{+}=\bar r_i+\kappa_r s_i^r,\qquad U_i^{-}=\bar u_i-\kappa_u s_i^u,\qquad S_i=U_i^{-}+\lambda_o\hat o_i-\lambda_r R_i^{+}
Equations (4–5): bars are ensemble means; \(s_i^r,s_i^u\) measure predictive dispersion. Frozen multipliers \(\kappa_r,\kappa_u\) form risk/utility bounds. Score \(S_i\) rewards utility and estimated opportunity \(\hat o_i\), and penalizes risk through nonnegative weights. e-verifier
Gi=gi1[Ri+τr]1[UiU0ϵu]1[o^iτo]1[Qiτq]1[SiS0τm],i>0.\begin{aligned}G_i={}&g_i\mathbf{1}[R_i^{+}\le\tau_r]\mathbf{1}[U_i^{-}\ge U_0^{-}-\epsilon_u]\\&\cdot\mathbf{1}[\hat o_i\ge\tau_o]\mathbf{1}[Q_i\ge\tau_q]\mathbf{1}[S_i-S_0\ge\tau_m],\quad i>0.\end{aligned}
Equation (6): indicators enforce risk threshold \(\tau_r\), tolerated utility loss \(\epsilon_u\), opportunity threshold \(\tau_o\), confidence threshold \(\tau_q\) and margin \(\tau_m\). \(Q_i\) is minimum satisfaction confidence over active events; subscript zero denotes the nominal candidate. \(G_i=1\) authorizes an override. e-gatese-verifier

5. Method in detail

5.1 Turn a coordination requirement into a noncompensable check

Reader analysis

Reader interpretation: think of a shared-object lift as a candidate that must satisfy several obligations at once. The paper's synchronization contract uses contact order, phase progress and bounded left–right delay; role compatibility considers object assignment and grasp ownership; collision convergence uses separation and relative approach. Task phase activates the relevant rows before outcomes are known. Equation (2) then multiplies their binary checks, so a failed required row makes the candidate inadmissible. This is the causal role of the contract: estimated task utility cannot buy its way past a violated obligation. Passing the predicates is only a prerequisite, because future-dependent ambiguity remains. The event-conditioned verifier interprets the proposed motion with the active obligation in view, supplying evidence that deterministic geometry and phase checks alone cannot settle. e-contracte-verifiere-interface

5.2 Distinguish an admissible alternative from a justified override

Source description

After admissibility, CoWAM still asks whether changing the nominal action is warranted. The ensemble raises the risk estimate and lowers the utility estimate using predictive dispersion and frozen calibration multipliers. The score rewards conservative utility and opportunity while penalizing risk, but ranking by this score is insufficient: every gate in Equation (6) must pass. The opportunity gate asks whether a useful alternative is predicted to exist; the margin gate requires a decisive improvement over the nominal score. Utility retention protects progress, while risk and event confidence constrain uncertain interventions. If no alternative qualifies, a valid nominal action is preserved; otherwise the contract fallback is invoked. The paper calls this abstention but does not specify every task's fallback action. Table 3 shows why the distinction matters: permissive variants gain positive validity while sharply increasing matched-negative errors. e-verifiere-gatese-traininge-ablation

5.3 Read the experiment as two complementary tests of selection

Reader analysis

Reader interpretation: the strongest design choice is fixing what each selector is allowed to choose. Identical ordered candidates and predicted futures prevent a better proposal generator from being mistaken for a better selector. Each decision is committed before shared oracle labeling, and candidate records within a pool are not independent statistical trials. The coordination audit then asks whether selection is valid on positive opportunities and restrained on matched negatives. Natural closed-loop evaluation asks a different question: do those decisions produce completed tasks over full episodes? The positive answers are 140/150 valid selections and 151/240 successful episodes, but their denominators are deliberately different. Finally, the 176/240 oracle result distinguishes remaining selection error from failures that the fixed proposal pool cannot solve. This interpretation stays within simulated tasks and the supplied evaluation protocol. e-audite-protocole-evente-natural

5.4 Training and inference

During training

Source description

The policy and WAM remain frozen. Table A2 specifies three verifier instances, a 60/20/20 train/validation/test split by task-seed group, AdamW at 3×10⁻⁴, batch 128 and 100 epochs. Appendix A identifies 1,200 test groups containing 9,600 candidate records. Calibration and normalization are frozen before evaluation. e-traininge-verifier

Reader analysis

The PDF does not specify the verifier backbone, complete loss functions or loss weights. It lists output targets and optimization settings, which are insufficient to reconstruct training exactly. e-verifiere-training

During inference

Reader analysis

Default settings include eight candidates, risk threshold 0.20, utility lower floor 0.55, selective margin 0.12, event confidence 0.70 and enabled abstention. The text does not fully map the tabulated absolute utility floor onto the nominal-relative inequality. e-traininge-gates

Source description

The online selector uses predicted evidence, not future simulator outcomes. Shared oracle execution labels candidates afterward for task success, coordination, progress and failure mode; oracle performance is an offline proposal ceiling. e-audit

5.5 Implementation flow

  1. Freeze the proposal interface

    At each replanning step, preserve candidate identities and order; index zero is nominal. The primary proposer is X-WAM. CoWAM neither resamples candidates nor updates the proposer. e-interfacee-environment

  2. Activate coordination obligations

    Task phase activates synchronization, role compatibility and collision-convergence predicates. Contact order, grasp ownership, separation and approach trends supply typed evidence. Every active obligation must pass; unrelated utility cannot compensate for a failure. e-contract

  3. Estimate future-dependent evidence

    Three independently seeded verifiers consume observation, action, predicted future and event type. Heads estimate satisfaction, risk, utility, opportunity and confidence. Ensemble dispersion makes risk more pessimistic and utility less optimistic. e-verifiere-training

  4. Qualify and execute

    An alternative must pass contract, risk, utility-retention, opportunity, event-confidence and score-margin gates. Choose the highest-scoring eligible alternative, breaking ties by proposer order. Otherwise preserve a contract-valid nominal candidate or use the frozen fallback, then replan. e-gatese-interface

6. Experiments & results

CoWAM uses typed coordination obligations and calibrated future evidence to decide whether an existing bimanual policy action should change. Its frozen proposer supplies alternatives; the selector cannot invent missing actions. Outcome-blind comparisons report better coordination selection and simulated closed-loop completion, while conservative gates limit unsupported overrides.

6.1 Read the original evidence

Table 1. Separate finding a valid candidate from avoiding an unsupported override. Original paper, p. 4 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Begin with the denominators in the header. Valid selection is measured on 150 oracle-confirmed opportunities; false and harmful intervention use 180 matched contract-valid negatives. They answer different questions and should not be combined into one success percentage. Compare the bold CoWAM row with Contract-only and the RGB-D selector, then inspect the event-family rows. Each family contributes 50 positive opportunities and 60 negatives, so the lower block reveals whether one event type dominates the aggregate. Policy Top-1 has zero intervention errors because it never overrides. The Oracle Upper Bound sees outcomes after commitment and is an offline reference. e-evente-protocole-audit

What it supports. CoWAM selects valid candidates on 140/150 opportunities versus 115/150 for Contract-only and 126/150 for the RGB-D selector. It records five false and one harmful intervention among 180 negatives. The family counts, 47, 46 and 47 of 50, support improvement across all three evaluated obligations.

Where the evidence stops. This is a constructed event audit, not the natural frequency of deployment opportunities. A harmful rate of 1/180 is conditional on this negative cohort; it is neither a universal failure probability nor a physical-robot safety estimate.

Table 2. Check whether better selection produces more completed simulated tasks. Original paper, p. 5 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Read panel (a) as an episode-level test: each method has 240 episodes from 30 held-out seeds on each of eight RoboTwin tasks. The success column uses strict simulator completion, while coordination validity and intervention errors remain separate metrics. Compare CoWAM with Selective Control, the strongest evaluated selective baseline in this block. Then read panel (b) row by row. The Success column is the additional number of completed episodes over Selective Control, rather than a percentage. Its total of 23 explains the aggregate improvement. The oracle column records remaining proposal headroom and is unavailable to the deployable selector. e-naturale-protocole-discrepancies

What it supports. CoWAM reaches 151/240 successes, or 62.9%, compared with 128/240, or 53.3%, for Selective Control. The reported 9.6-percentage-point gain is positive on every task, ranging from one additional Scan Object success to four additional successes on each of two tasks. The oracle still reaches 176/240.

Where the evidence stops. These are held-out seeds of evaluated simulated task definitions. The table supplies no confidence intervals or physical deployment results. Appendix C's claim of two-to-four extra successes per task conflicts with Scan Object's visible one-episode gain.

Figure 6. Use the camera sequence to connect the reported task outcome to visible object transport. Original paper, p. 6 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Read left to right within each row, following initial, middle and terminal states. The first two rows show the nominal failure from front and head cameras. The lower two rows show the corresponding CoWAM execution, again from front and head views. The head camera makes the can and yellow basket easy to track; the front view supplies arm and workspace context. At the terminal stage the CoWAM sequence places the can inside the basket. Interpret these as author-selected execution images, not the verifier's predicted future or a complete record of every intermediate contact. e-baskete-naturale-gates

What it supports. The source presents this case as successful object acquisition, transport and basket placement under CoWAM, contrasted with nominal failure. The visible terminal placement helps explain what task completion means in this example. It complements the aggregate experiment but does not establish how frequently the same intervention succeeds.

Where the evidence stops. This isolated case provides neither gate values nor the selected candidate's full decision record. Camera endpoints cannot establish every timing, contact or collision obligation, and a chosen success sequence is not evidence of general robustness.

6.2 Results and evaluation conditions

Task & protocolReported resultComparison & interpretation
Natural closed-loop success across eight RoboTwin 2.0 tasks

30 paired held-out seeds per task; 240 simulated episodes per method; strict simulator completion within a fixed horizon.

151/240 successes (62.9%); 95.4% coordination validity; 3.3% false and 0.4% harmful intervention.

Task success; coordination validity; false/harmful intervention rates

Selective Control: 128/240 (53.3%); Policy Top-1: 96/240 (40.0%); offline oracle: 176/240 (73.3%).

The reported gain over Selective Control is 9.6 percentage points, with 32 versus 9 discordant pairs and reported exact paired p=4.3×10⁻⁴. Each task improves by 1–4 episodes; physical-robot transfer is untested. e-naturale-protocol

Coordination-valid selection at event opportunities

180 event-stress clusters across eight tasks; 150 oracle-confirmed opportunities and 180 matched contract-valid negatives.

140/150 valid (93.3%); 5/180 false (2.8%); 1/180 harmful (0.6%).

Valid selection on positives; false/harmful intervention on negatives

Contract-only: 115/150, 20/180 and 3/180; RGB-D selector: 126/150, 10/180 and 2/180.

The 16.7-point gain over Contract-only has 27 versus 2 discordant pairs (reported p=1.6×10⁻⁶). Family counts are 47/50, 46/50 and 47/50. These opportunity-conditioned rates are distinct from episode success. e-evente-protocol

Conservative-gate ablation

The same 150 positive and 180 matched-negative units as the coordination audit.

Full: 93.3% / 2.8% / 0.6%.

Valid / false / harmful rates

No uncertainty bound: 94.7% / 14.4% / 3.3%; no baseline preservation: 96.7% / 28.9% / 7.8%.

Higher positive selection alone favors permissive variants; matched-negative errors expose the cost. Contract-only removes both learned evidence and gates, so it is not an isolated test of event conditioning. e-ablatione-contract

Learned coordination-risk ranking

1,200 task-seed-disjoint test groups; 9,600 correlated candidate records.

0.92; 0.89; 0.03.

AUPRC; pair accuracy; expected calibration error

Scalar verifier: 0.80 / 0.79 / 0.06; shuffled futures: 0.63 / 0.62 / 0.14.

Candidate-specific future evidence supports discrimination and calibration in this test distribution; it is not a bound on deployment risk. e-rankinge-training

Risk at selective coverage

Retained subsets of the 1,200 ranking groups, reused rather than independent experiments.

25% coverage: 2/300 (0.7%); full coverage: 70/1,200 (5.8%).

Coordination-risk violation count/rate

Future-Consensus: 10/300 (3.3%) and 269/1,200 (22.4%).

Risk increases as more groups are retained; these rates use retained-group denominators, not the matched-negative denominator. e-coverage

Candidate-count and proposer robustness

80 states per candidate count; separately, 120 paired pools per proposer regime over six tasks.

K=4: 26/28 rescues; K=32: 38/40. X-WAM, LeWorldModel and mixed-pool success: 78/120, 75/120 and 85/120.

Rescues per available opportunity; proposer-conditioned task success

Corresponding proposer baselines: 66/120, 60/120 and 70/120.

Reported success gains are 10.0, 12.5 and 12.5 percentage points. Reuse of the selector interface does not establish unrestricted transfer to new embodiments. e-robustness

6.3 Ablations and diagnostic examples

Read component removals and qualitative examples within their stated evaluation conditions.

Table 3. Read the apparent gains from permissive selection alongside their intervention cost. Original paper, p. 6 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Start with panel (a), which reuses the 150 positive and 180 negative audit units. Removing uncertainty bounds or baseline preservation raises positive validity, but examine the adjacent false and harmful columns before judging that improvement. Contract-only removes both the learned verifier and the full gate stack, so its comparison measures a bundle of components. Panel (b) changes the evidence representation and uses a separate ranking block of 1,200 test groups containing 9,600 candidate records. Higher AUPRC and pair accuracy indicate stronger discrimination; lower ECE indicates lower measured calibration error. Those quantities are not episode-success rates and cannot be compared numerically with panel (a)'s percentages. e-ablatione-rankinge-traininge-discrepancies

What it supports. Without uncertainty bounds, validity rises from 93.3% to 94.7%, while false intervention rises from 2.8% to 14.4% and harm from 0.6% to 3.3%. Removing baseline preservation is still more permissive. Separately, shuffled futures reduce AUPRC from 0.92 to 0.63, supporting the relevance of candidate-specific temporal evidence.

Where the evidence stops. Table 3's no-event-conditioning row reports 74.7% validity, whereas Table A11's RGB-D + future, no event row reports 128/150. The PDF does not reconcile these configurations. Do not treat them as a single precisely isolated effect.

Table A14. Calibration supports an operating tradeoff whose denominator changes with coverage. Original paper, p. 14 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Read the retained-set columns first. Coverage grows from 300 to 1,200 groups, so each rate divides its violation count by that row's retained population. For example, CoWAM's two violations at 25% coverage become 0.7% after rounding; its 70 violations at full coverage become 5.8%. Compare methods horizontally at the same coverage, then follow each method downward as more groups are retained. These are reused groups from the learned-ranking evaluation. The table therefore examines an empirical selection–risk tradeoff without adding a new independent experiment or measuring the matched-negative harmful-intervention rate. e-coveragee-ranking

What it supports. CoWAM has lower reported coordination-risk violation rates at every shown coverage: 0.7% versus 3.3% at 25%, and 5.8% versus 22.4% at full coverage. The rising CoWAM rate makes the cost of broader coverage visible, even though its curve remains below Future-Consensus throughout the evaluated range.

Where the evidence stops. The table is an observed risk–coverage relationship on the evaluated groups. It does not prove calibrated bounds under a distribution shift, and its 5.8% full-coverage rate must not be substituted for closed-loop harm or event-audit error.

7. Analysis & limitations

7.1 What the evidence leaves open

Reader analysis

Evidence is from RoboTwin 2.0 and held-out seeds of the evaluated task definitions. New-task and physical-robot generalization remain open. The oracle–CoWAM success gap also shows residual selection headroom within the fixed proposal pool. e-protocole-natural

Reader analysis

An unresolved ablation discrepancy remains: Table 3 lists no event conditioning at 74.7% valid, 10.0% false and 1.7% harmful; Table A11 lists RGB-D + future, no event at 128/150, 10/180 and 2/180. The source does not explain the configuration difference. Appendix C also says per-task gains are 2–4, while Tables 2/A5 show Scan Object improves by one. e-discrepancies

Reader analysis

Conservative bounds are empirical ensemble/calibration constructions; the PDF supplies no formal deployment guarantee. Exact predicates, fallback actions, calibration coefficients and several gate settings remain unspecified in the supplied paper. e-contracte-verifiere-training

7.2 Questions for discussion

  1. How much of the gate benefit survives when variants are compared at equal override frequency?
  2. Can frozen calibration maintain matched-negative precision under a shift in contact timing or object geometry?
  3. What distinct configurations account for the two no-event ablation results?

8. Reproducibility audit

8.1 Requirements and known gaps

Author claim

Reproduction requires pinned simulator/model revisions, frozen task-seed matrices, ordered candidate-pool digests, committed outcome-blind decisions, shared oracle labels and paired aggregation. The paper says its submitted archive contains these artifacts; their availability is not established by that statement. e-artifacts

Reader analysis

Reported selector cost is 88 ms and 4.8 GB on one RTX 5880 Ada, versus 41 ms/3.1 GB for Policy Top-1; offline oracle rollout costs 2,460 ms. Training duration and a complete training-data inventory are not specified. e-runtimee-training

8.2 Proposed reproduction checks

The following checks are proposals motivated by the paper. They have not been run as part of this reading.

Check 1: Test gate benefits at a matched override frequency

Reader-proposed experiment: regenerate the fixed event pools and compare full CoWAM, point-estimate uncertainty and no-baseline-preservation variants with identical verifier outputs, candidate order and oracle labels. First reproduce their published operating points. Then tune only validation-set thresholds to match the full method's override frequency, freeze them, and evaluate the same held-out clusters. Report valid/150, false/180, harmful/180, abstentions and paired differences. If the error advantage disappears after matching override frequency, much of the apparent gate benefit reflects conservatism; if it persists with comparable valid selection, the gates discriminate risky interventions beyond merely reducing their number. No such reproduction was performed for this reading. e-gatese-ablatione-protocol

Check 2: Separate future alignment from event conditioning

Reader-proposed experiment: resolve the configurations behind Tables 3 and A11 from their manifests, then compare a controlled two-by-two design: aligned versus within-pool shuffled predicted futures, each with versus without event conditioning. Keep action identities, inputs other than the tested factors, data splits, capacity and optimization budget fixed; calibrate each variant on validation groups only. Commit decisions before shared labeling and report ranking, calibration and audit errors on the same held-out units. A future-alignment effect should reduce discrimination when futures are shuffled; an event-conditioning effect should survive the controlled comparison. Failure to recover either would weaken that mechanism claim. Report both original no-event settings separately if they differ. This is a proposed check, not an executed result. e-rankinge-discrepanciese-traininge-audit

8.3 Reading coverage

Visual audit: Visually inspected the title/author/version page, Figure 2 architecture, formal equations on p. 3, main coordination and closed-loop tables, Table 3 ablations, qualitative Figures 5–6, Appendix A configuration/contract tables, Table A5 task counts, and Tables A10–A15 including the conflicting no-event row and risk–coverage counts. All six final original-PDF crops were viewed after extraction. Other PDF pages were read completely as supplied text, including their captions, but their images were outside this visual pass. Separate media, code and the claimed reproducibility archive were not supplied or inspected.

PDF pages inspected for this edition: 1, 2, 3, 4, 5, 6, 10, 12, 14. Appendix coverage: reviewed.

Original figures and tables remain the work of the source’s authors. Extractions preserve their scientific content; any HTML wrapper layout is disclosed with each figure. The surrounding reading notes are our own.

Text reading scope & known omissions
  • PDF p. 1: title, authors, affiliations, arXiv version stamp, Abstract and Introduction
  • PDF pp. 2–3: Related Work; Methods, including candidate interface, contracts, verifier, calibration and selective intervention
  • PDF pp. 4–7: Experiments, Setup, Main Results, Ablations, Robustness Analysis, Qualitative Cases, Discussion and Limitations, Conclusion
  • PDF p. 8: References
  • PDF pp. 9–11: Appendix A, Additional Method Details; B, Experimental Protocol; C, Complete Result Decompositions; D, Robustness, Ablations, and Runtime
  • PDF pp. 11–16: Appendix E, Reproducibility and Evaluation Coverage; Tables A5–A16 and associated result text
  • PDF pp. 15–17: Appendix G, Qualitative Case Supplement and Figures A7–A9 captions; no Appendix F heading appears in the supplied text

Outside the original text pass

  • Text extraction does not reconstruct figure images; inspect the retained PDF for figures and equation/table layout.
  • Separate supplemental material availability has not been fully verified.
  • The extraction limitation was addressed by visual inspection of PDF pages 1, 2, 3, 4, 5, 6, 10, 12 and 14 and all six final crops. Other pages were read as complete extracted text; their figure images were not visually inspected.
  • The submitted reproducibility archive, code, weights and separate media were not supplied or inspected; no experiments were reproduced.
  • Identity matches the catalog: the observed title and all eight authors agree. The inspected artifact is arXiv:2608.02578v1 [cs.RO], 3 August 2026. No other revision or edition was supplied or compared.

The visual audit above records the subsequent illustrated pass.

8.4 Traceable evidence

e-identityPDF p. 1, title/author block and left-margin arXiv stampInspect

Title matches the supplied observed title. Authors are Shuaijun Liu, Qifu Wen, Shuyang Hao, Qi Luo, Chenglong Zhang, Feiyang You, Chengyu Wu and Ningxin Su. Affiliations are HKUST (Guangzhou), Boston University and Shanghai Jiao Tong University. Stamp: arXiv:2608.02578v1 [cs.RO], 3 Aug 2026.

Go to primary source ↓
e-problemPDF p. 1, Abstract and Introduction; Figure 1Inspect

Bimanual timing, arm roles, spatial compatibility and phase consistency motivate conservative intervention in nominal policy actions.

Go to primary source ↓
e-interfacePDF pp. 2–3, Figure 2; World-Action Candidate Interface, Eq. (1); p. 9, Decision ProcedureInspect

Frozen ordered action–future pools retain nominal index zero and candidate identities. CoWAM preserves, overrides or abstains without changing the proposer; execution is followed by replanning.

Go to primary source ↓
e-contractPDF p. 3, Coordination Contracts, Eq. (2); p. 9, Contract Structure; p. 10, Table A1Inspect

All active synchronization, role and collision predicates must pass. Contract state includes event activation, evidence requirements, thresholds and fallback; inactive events impose no constraint.

Go to primary source ↓
e-verifierPDF p. 3, Event-Conditioned Contract Evidence, Eqs. (3–5)Inspect

Event-conditioned outputs estimate satisfaction, risk, utility, opportunity and confidence. Ensemble means and dispersion form conservative risk/utility bounds; the frozen score combines utility, opportunity and risk.

Go to primary source ↓
e-gatesPDF p. 3, Selective Policy Intervention, Eq. (6); p. 9, Decision ProcedureInspect

Contract, risk, utility retention, opportunity, event confidence and margin jointly qualify an alternative. Highest score wins, ties follow proposer order; otherwise preserve a valid nominal action or invoke fallback.

Go to primary source ↓
e-trainingPDF p. 9, Verifier and Calibration; p. 10, Table A2(a–b)Inspect

Four predicted future steps, three independently seeded verifiers, 60/20/20 task-seed group split, AdamW 3×10⁻⁴, batch 128 and 100 epochs are specified. Defaults are K=8, risk 0.20, utility floor 0.55, margin 0.12, confidence 0.70 and enabled abstention; 1,200 groups/9,600 records are identified as the test block.

Go to primary source ↓
e-auditPDF p. 3, final Methods paragraphs; pp. 9–11, Appendix B, Outcome-Blind PairingInspect

Selectors commit identical-pool decisions before shared simulator labeling. Beneficial changes repair a nominal failure without losing a required outcome; harmful changes lose an outcome; false changes lack task/coordination support. Oracle labels are offline.

Go to primary source ↓
e-protocolPDF p. 4, Experiments—Setup; p. 10, Table A3; p. 11, Metrics and Statistical Tests; p. 16, Table A16Inspect

Eight RoboTwin tasks use 30 natural seeds each. Event audit has 180 clusters, 150 opportunities and 180 matched negatives. Exact paired tests use clusters/episodes, not candidates. Ablations reuse pools; the ledger distinguishes 2,500 independent scientific units from candidate records and timing repetitions.

Go to primary source ↓
e-eventPDF p. 4, Table 1, CoWAM/Contract-only/RGB-D rows and event-family block; p. 12, Table A6; p. 13, Table A7Inspect

CoWAM records 140/150 valid, 5/180 false and 1/180 harmful; Contract-only 115/150, 20/180, 3/180; RGB-D 126/150, 10/180, 2/180. Family conversions are 47,46,47 of 50. Reported paired discordances are 27 versus 2, p=1.6×10⁻⁶.

Go to primary source ↓
e-naturalPDF p. 5, Table 2(a–b) and Main Results; p. 12, Table A5Inspect

CoWAM 151/240 successes versus Selective Control 128/240, Policy Top-1 96/240 and oracle 176/240. CoWAM validity/false/harm rates are 95.4%/3.3%/0.4%. Paired discordances 32 versus 9 give reported p=4.3×10⁻⁴. Scan Object rises 17→18; all eight tasks improve.

Go to primary source ↓
e-ablationPDF p. 6, Table 3(a) and Ablations discussionInspect

Full CoWAM valid/false/harm rates are 93.3/2.8/0.6%; removing uncertainty gives 94.7/14.4/3.3%; removing baseline preservation gives 96.7/28.9/7.8%. Contract-only removes learned evidence and gate stack.

Go to primary source ↓
e-rankingPDF p. 6, Table 3(b); p. 14, Table A10Inspect

CoWAM AUPRC/AUROC/pair accuracy/ECE/Brier are 0.92/0.96/0.89/0.03/0.07. Scalar AUPRC/pair/ECE are 0.80/0.79/0.06; shuffled futures 0.63/0.62/0.14; current-plus-action AUPRC is 0.75.

Go to primary source ↓
e-coveragePDF p. 14, Table A14; p. 11, Appendix E, Denominator LedgerInspect

At 25/50/75/100% coverage, retained N is 300/600/900/1,200; CoWAM violations are 2/7/24/70, reported 0.7/1.2/2.7/5.8%. Future-Consensus has 10/47/131/269. Coverage rows reuse ranking groups.

Go to primary source ↓
e-robustnessPDF p. 7, Table 4; p. 13, Tables A8–A9; p. 9, Appendix B, Tasks and Evidence AllocationInspect

K=4/8/16/32 uses 80 states each, opportunities 28/32/36/40 and CoWAM rescues 26/30/34/38. Proposer success is 78/120,75/120,85/120 versus 66/120,60/120,70/120 across X-WAM, LeWorldModel and mixed pools; transfer uses six tasks and 20 seeds per regime.

Go to primary source ↓
e-discrepanciesPDF p. 6, Table 3(a), No event conditioning; p. 14, Table A11, RGB-D + future, no event; p. 11, Appendix C, Complete Selector Counts; pp. 5/12, Tables 2/A5, Scan ObjectInspect

The no-event rows report 74.7%/10.0%/1.7% versus 128/150,10/180,2/180 without an explicit reconciliation. Appendix C states per-task gains of two to four, but Scan Object improves from 17/30 to 18/30.

Go to primary source ↓
e-environmentPDF p. 14, Table A13Inspect

Primary proposer is X-WAM; additional proposers are LeWorldModel and mixed pools. The host has eight RTX 5880 Ada GPUs, with one explicitly pinned GPU per job.

Go to primary source ↓
e-runtimePDF p. 14, Table A15(a–b)Inspect

Per-selection CoWAM latency/memory are 88 ms/4.8 GB, Policy Top-1 41 ms/3.1 GB; offline oracle rollout is 2,460 ms. Failure labels use 240 natural episodes and are non-exclusive.

Go to primary source ↓
e-artifactsPDF pp. 11, 14–15, Appendix E, Runtime Environment and Artifacts; Table A13 caption; Claim-Reproduction OrderInspect

The source states a submitted archive contains configurations, revision/checksum references, manifests, decisions, aggregates, statistical scripts and media. It prescribes frozen pools, pre-oracle commitment, shared labeling and paired aggregation.

Go to primary source ↓
e-basketPDF p. 6, Figure 6 and captionInspect

Front/head views at initial, middle and terminal stages contrast nominal Place Can in Basket failure with author-labeled successful CoWAM acquisition, transport and placement.

Go to primary source ↓

8.5 Primary sources

Scroll across the image to inspect details. Press Esc to close.