PAPER REPORTENAll readings ↗

A Watermark for Vision-Language-Action and World Action Models

English reading report: Method, equations, original figures, experiments and reproducibility.

Authors: Yule Liu; Shuai Liu; Jiaheng Wei; Xinlei He

Affiliations: Hong Kong University of Science and Technology (Guangzhou); Wuhan University; Xi’an Jiao Tong University

Source: 2606.23574 ↗ · Catalog record

Reading: 151 / 558 · 6 original figures & tables · ~20 min ·

1. Paper overview

In one sentence: A keyed sampler leaves recoverable provenance in executed actions, but the evidence depends on calibration and continued use of that sampling process. e-threate-injectione-mape-identificatione-distillation

At a glanceWhat to know
Research problem
Source description

Can an owner identify a resold black-box robot service from its final actuator commands, while keeping the policy useful? The protected service must still invoke the keyed sampler. The owner has authorized audit access and its own base model; ordinary adversaries can wrap or post-process outputs but cannot replace the sealed sampler. e-threat

Core mechanism
Source description

A sampling-time fingerprint shared by continuous-noise VLA and WAM generators, requiring no watermark training or new policy architecture. e-architecturee-injection

A key reported resultClean ownership verification across four model–suite cells: All four group AUCs and TPRs are 1.000. Per-episode AUC: LingBot 1.000/1.000 and π0.5 0.901/0.852 (LIBERO/RoboTwin).

Per-episode AUC; group AUC and TPR at nominal 1% FPR; task success rate (SR).. π0.5 and LingBot-VA on LIBERO-10 and RoboTwin; partial executed-action observation, MAP, no attack; 16 bootstrapped episodes per decision.

Plain→marked SR: LingBot 0.94→0.96 and 0.57→0.53; π0.5 0.96→0.96 and 0.50→0.57. Aggregation helps π0.5 substantially. Utility shifts span −4 to +7 percentage points; these estimates lack uncertainty intervals in Table 1. Nominal FPR is not uniformly realized. e-maine-protocole-calibration

Reading caution
Reader analysis

The binary threshold calibrated at nominal 1% yields plain-rollout FPR 2.0% with one episode and 2.4% with sixteen for LingBot/RoboTwin. Good AUC cannot establish correct fixed-threshold false-attribution control. e-calibration

Core contributions

  • Source description

    A sampling-time fingerprint shared by continuous-noise VLA and WAM generators, requiring no watermark training or new policy architecture. e-architecturee-injection

  • Source description

    A partial-action MAP verifier with decoy-key calibration, rollout aggregation and multi-key attribution, evaluated against output attacks, owner variants and stronger boundary stress tests. e-mape-scoree-identificatione-variantse-distillation

Figure 1. A watermark enters at sampling and is tested later through partial actions. Original paper, p. 2 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Start at the left: observations and the instruction condition the existing VLA/WAM, while the secret key influences its initial sampling noise. The plus symbol is schematic; Equation (1) specifies a weighted Gaussian mixture, not unrestricted addition to actions. Follow the raw output grid to the smaller partial-observation block, then MAP optimization and candidate-key comparison. The stacked dashed boxes indicate repeated rollouts. At the right, aggregated scores support either a yes/no test or selection of a matching key. Read the figure together with Equations (2–3): recovery must reproduce the visible channels through the owner's generator and the modeled post-processing. e-architecturee-injectione-mape-threat

What it supports. The common interface is the differentiable path from seed to action. This lets the owner apply one verification strategy to direct VLA generation and future-scene-mediated WAM generation. The action grid is an audit surface; the diagram does not propose a new policy or show watermark training.

Where the evidence stops. The diagram omits the recorded public nonce, keyed chunk schedule and details of post-processing. It does not establish that the audit can reconstruct missing context or an unknown controller transformation; these remain prerequisites or implementation questions.

2. Motivation

2.1 The problem and the proposed response

Source description

Can an owner identify a resold black-box robot service from its final actuator commands, while keeping the policy useful? The protected service must still invoke the keyed sampler. The owner has authorized audit access and its own base model; ordinary adversaries can wrap or post-process outputs but cannot replace the sealed sampler. e-threat

2.2 What this reading follows

A robot policy can hide its weights while exposing the motor commands that make it commercially useful. This paper asks whether those commands also reveal a protected sampling process. The owner marks selected initial noise vectors, then uses its own model to reconstruct plausible seeds from the suspect's partial action stream. Matching recovered seeds to secret references turns robot rollouts into provenance evidence. The reading hinges on three distinctions: fitting visible actions versus guessing hidden channels, detecting a key versus naming it with controlled false alarms, and surviving a service wrapper versus surviving a newly trained student. The original tables and diagnostics make each boundary concrete. e-threate-injectione-mape-identificatione-distillation

3. Research context

We place the paper in the collection through its world–action interface. The catalog labels and the reading’s assessment are shown separately.

Catalog dimensionRecorded classification
Major categoryFoundational work
ArchitectureNot applicable
Prediction paradigmNot applicable
QuadrantNot applicable

This table preserves the labels recorded at reading time. The current major category is Related resources. View the current classification.

3.1 Evidence-based assessment

Supports the recorded classification

Reader analysis

The recorded foundational evaluation/protocol classification fits this provenance method. It operates on existing VLA and WAM generators rather than proposing its own world-action architecture. One Model/Two Models and prediction-quadrant labels are therefore not applicable to the contribution, even though the host WAM abstraction predicts a future scene before actions. e-architecturee-threat

This is the collection’s architectural analysis, not a new related-work survey. Benchmark comparisons and their protocols appear in Section 6.

4. Problem formulation

4.1 Inputs and outputs

InputsOutputs
  • Policy observations and instruction o_c; owner generator F_θ; Gaussian seed z_c.
  • Secret candidate key, recorded public episode nonce and chunk positions; observed executed actions y_c.
  • Deployment: action chunks from the existing policy.
  • Audit: calibrated key scores, binary verification, or a ranked key with optional open-set rejection.

4.2 Equations and their role

zcfp=1β2zc+βrcz_c^{\mathrm{fp}}=\sqrt{1-\beta^2}\,z_c+\beta r_c
Equation (1): z_c is base Gaussian noise, r_c the keyed reference and β injection strength. Under the stated independent standard-Gaussian assumptions, the mixture has identity covariance marginally over fresh references. At the deployed β=1, the selected seed equals r_c. e-injectione-hyperparameters
z^c=argminzg(PC(Fθ(z;oc)))yc222σobs2+λz2z22\hat z_c=\arg\min_z\frac{\|g(P_C(F_\theta(z;o_c)))-y_c\|_2^2}{2\sigma_{\mathrm{obs}}^2}+\frac{\lambda_z}{2}\|z\|_2^2
Equation (3): F_θ generates actions conditioned on o_c; P_C selects visible channels, g models post-processing, and y_c is the observed chunk. σ_obs weights action fitting and λ_z weights the latent prior. The objective recovers a plausible explanatory latent; it does not guarantee unique recovery. e-map
Ze(k)=se(k)μe(k)σe(k)+εZ,TG(k)=eGZe(k)Z_e(k)=\frac{s_e(k)-\mu_e^-(k)}{\sigma_e^-(k)+\varepsilon_Z},\qquad T_G(k)=\sum_{e\in G}Z_e(k)
Equations (6–7): s_e is the matched-filter score for key k in episode e; μ_e^- and σ_e^- are decoy mean and scale, and ε_Z prevents division by zero. G is the rollout group. Independence and a Gaussian approximation underlie the proposed square-root budget scaling. e-score

5. Method in detail

5.1 Why a seed can carry a mark without a new policy

Source description

The proposed intervention sits before the existing action generator. For each episode and chunk, the key and recorded public nonce determine both a reference and a selection schedule. Selected chunks use Equation (1); unselected chunks retain ordinary noise. Under the paper's independent standard-Gaussian assumptions, the weighted mixture preserves the marginal seed distribution while correlating it with a reference the owner can reproduce. With the experimental β=1, a selected seed is entirely the keyed reference. The VLA then generates actions directly; the WAM passes through a predicted future-scene representation. No policy weights need to learn a secret trigger. This explanation is conditional on the stated reference law: the experiment's separate description of four band-limited Gaussian tones is not reconciled with the coordinate-wise i.i.d. construction in the supplied text. e-injectione-architecturee-hyperparameterse-protocol

5.2 Why recovery fits actions rather than reversing an incomplete endpoint

Source description

The audit sees the commands that actually reach the robot. A raw action chunk may contain channels that were never executed, and the controller can change those that remain. Reverse ODE recovery needs a full endpoint, making hidden-channel completion consequential. The MAP objective instead generates a candidate action chunk, projects it onto visible channels, applies the modeled transformation and compares it with the recorded command. A latent prior discourages large explanatory seeds, while multiple starts address the optimization landscape. The owner retains the lowest-loss candidate and tests its correlation with keyed references. Table 2 shows the precise advantage: partial MAP remains useful without the training padding convention, although correctly padded ODE is stronger. Thus the result supports a verifier with fewer hidden-endpoint assumptions, not guaranteed recovery of the exact original seed. e-mape-recoverye-threat

5.3 Why a strong group score is only one part of attribution

Reader analysis

After recovery, decoy keys estimate how much apparent alignment an episode produces for the wrong key. Standardized scores are summed across rollouts, and the paper's rate law predicts improved separation when independent episodes contain a positive signal. My reading is that this creates three separate obligations. The owner must show detection power, demonstrate error control on genuinely unmarked policies, and specify which transformations preserve the protected process. Table 1 addresses power; Figure 4 exposes a mismatch between decoys and the plain-policy null. Open-set identification adds a gallery-wide rejection threshold calibrated from plain impostors. Finally, Figure 9 shows why the scope matters: a fresh student can reproduce behavior without inheriting the keyed sampler. Neither high clean AUC nor a large key space resolves those calibration and process-replacement boundaries. e-scoree-maine-calibratione-identificatione-distillatione-security

5.4 Training and inference

During training

Source description

Injection leaves existing policy weights unchanged. MAP updates the input latent at audit time, not the policy. LoRA fine-tuning and compression are separate descendant tests. e-architecturee-mape-variants

Source description

For the π0.5 distillation stress test, teacher actions relabel fixed observations and an attention-only LoRA student trains for 1,500 steps, then runs without the keyed sampler. e-distillation-setup

During inference

Source description

Main settings use β=1, m=5 and 32 decoys. P is 1 for π0.5, 6 for LingBot/LIBERO-10 and 2 for LingBot/RoboTwin. MAP uses λ_z=1 and σ_obs=10^{-4} for π0.5 or 10^{-3} for LingBot. These are audit settings, not a newly learned controller. e-protocole-hyperparameters

5.5 Implementation flow

  1. Follow the existing seed-to-action path

    A VLA maps noise directly to actions. The paper's WAM abstraction first predicts a future-scene representation h, then maps h to actions. Both paths are treated as a differentiable F_θ; the watermark does not introduce a new planning loop. e-architecture

  2. Select and mark chunks

    A key-and-nonce selector marks at most m chunks with maximum gap P. Selected seeds mix base noise with the keyed Gaussian reference; unselected seeds remain ordinary. Fresh references are essential to the stated marginal noise-law argument. e-injection

  3. Fit only the audit view

    Project generated actions onto executed channels, apply the modeled post-processing, and optimize the latent against observed commands with a Gaussian prior. Try several random starts and keep the lowest-loss fit. Recovery uses the owner's base generator even for a suspect descendant. e-threate-map

  4. Calibrate and decide

    Match recovered latents with each key's references, standardize against decoys, and sum episode scores. A bounded global lag search addresses delay. Binary thresholds use grouped decoy scores; open-set identification instead calibrates the gallery maximum on plain impostor rollouts. e-scoree-identification

6. Experiments & results

Keyed latent-provenance verification marks selected sampling seeds in an existing robot policy, then recovers key evidence from executed action channels using the owner's differentiable generator. It supports both direct VLA and future-scene-mediated WAM policies. Sixteen-rollout groups separate the carried key in the clean tests, but calibration drift, strong output processing and replacement by a distilled student limit the ownership claim.

6.1 Read the original evidence

Table 1. Aggregation produces perfect reported clean separation, with variable utility changes. Original paper, p. 7 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Read each row as one model–suite setting. The utility block compares the same policy without injection, pl, against its fingerprinted version, fp. SR is a fraction, so a change of 0.07 is seven percentage points. Steps are mean episode length, not audit optimization time. In the clean and robust blocks, AUC refers to one episode, while AUC with subscript 16 and TPR use sixteen-episode groups. The robust columns average canonical-strength output attacks. Follow the π0.5 rows horizontally: aggregation improves separation even when the individual-episode AUC is materially below the LingBot values. e-maine-protocole-calibratione-attack-table

What it supports. Every clean setting reports group AUC and nominal-1%-FPR TPR of 1.000. Utility is less uniform: LingBot/RoboTwin drops from 0.57 to 0.53 SR, while π0.5/RoboTwin rises from 0.50 to 0.57. The lowest canonical attack-average TPR is 0.840 for π0.5/LIBERO-10.

Where the evidence stops. The robust block is an average, not the weakest attack. Utility has no uncertainty intervals here. Figure 4 also shows that nominal 1% FPR becomes 2.4% on plain LingBot/RoboTwin groups of sixteen.

Figure 4. Decoy-key calibration does not always control errors on unmarked policies. Original paper, p. 8 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. The vertical axis is an error rate in percent, so lower values are preferable here. The dashed horizontal line is the nominal one-percent target used when setting thresholds from decoy keys. Blue circles evaluate single episodes; orange squares evaluate sixteen-episode groups. Every point uses plain rollouts with no injection, making this a check of the actual unmarked-policy null rather than just another wrong-key score. Compare each pair with the dashed target before comparing models. The LingBot/RoboTwin pair is the exception: both points exceed the target, and grouping does not repair that calibration mismatch. e-calibratione-scoree-protocol

What it supports. LingBot/RoboTwin yields realized false positives of 2.0% for one episode and 2.4% for sixteen, despite a nominal 1% threshold. The other three settings are at or below the target. Strong detection and correct attribution-error calibration are separate empirical properties.

Where the evidence stops. The plotted zero values are finite-sample observations, not proof of zero false-attribution probability. No error bars are shown. The paper also warns that post-processing, especially jitter, can aggravate the calibration drift.

6.2 Results and evaluation conditions

Task & protocolReported resultComparison & interpretation
Clean ownership verification across four model–suite cells

π0.5 and LingBot-VA on LIBERO-10 and RoboTwin; partial executed-action observation, MAP, no attack; 16 bootstrapped episodes per decision.

All four group AUCs and TPRs are 1.000. Per-episode AUC: LingBot 1.000/1.000 and π0.5 0.901/0.852 (LIBERO/RoboTwin).

Per-episode AUC; group AUC and TPR at nominal 1% FPR; task success rate (SR).

Plain→marked SR: LingBot 0.94→0.96 and 0.57→0.53; π0.5 0.96→0.96 and 0.50→0.57.

Aggregation helps π0.5 substantially. Utility shifts span −4 to +7 percentage points; these estimates lack uncertainty intervals in Table 1. Nominal FPR is not uniformly realized. e-maine-protocole-calibration

Latent recovery under hidden-channel uncertainty

Same 50 marked and 50 plain π0.5/LIBERO-10 rollouts; clean base policy; 7 of 32 channels observed.

Partial MAP 0.775; full-observation MAP 0.994.

Per-chunk separation AUC.

Partial ODE: no pad 0.492, correct zero pad 0.969, wrong unit pad 0.678, secret random pad 0.508; full ODE 0.809.

MAP avoids dependence on an unknown completion convention. Correctly informed ODE beats partial MAP, so this is not a universal MAP superiority result. e-recovery

Clean identification in a 33-key gallery

One carried key and 32 decoys; closed-set ranking and open-set rejection calibrated with plain impostors.

At 16 episodes, rank-1 is 1.00 throughout; DIR is 1.00 except π0.5/RoboTwin at 0.99.

Rank-1 accuracy; DIR at 1% FAR.

Single-episode rank-1 is 0.44 and 0.23 for π0.5/LIBERO and RoboTwin, versus 1.00 in both LingBot cells.

Identification requires winning the gallery and, for DIR, clearing rejection. These results concern this gallery size. e-identification

Synchronization-search ablation

Table 7, π0.5/LIBERO-10; fixed delays; identical lag-zero versus global-search comparison.

At 16 episodes, search gives 1.00/0.96/0.53 for delays 1/2/3.

TPR at nominal 1% FPR.

Lag-zero gives 0.28/0.02/0.00. With 64 episodes, search gives 1.00 throughout.

Alignment is useful in this ablation, but Table 6 reports different delay values without explaining a protocol difference. e-synce-attack-table

Output and descendant robustness boundaries

Table 6, partial MAP audit; output processing applied to marked and plain pools; owner variants recovered with the base model.

At 64 episodes, π0.5/LIBERO EMA-lo and jitter-hi reach only 0.09 and 0.17. At 16 episodes, π0.5/RoboTwin action-side pruning gives 0.56 and int8 gives 0.94.

TPR at nominal 1% FPR.

LoRA variants reach 1.00 in every cell at 16 episodes. Table 1's weakest canonical attack-average TPR is 0.840.

An attack average is not a worst-case guarantee. Strong smoothing/jitter are reported to preserve utility, and Figure 10's smoothing curve conflicts with the low table value. e-maine-variantse-attack-tablee-budget-conflict

Behavior-cloning replacement of the protected sampler

LIBERO-10 student stress tests, evaluated without injection against a clean-student null.

Deployed zero-mean mark: 0.504 for π0.5 and 0.508 for LingBot.

Cross-student detection AUC.

Diagnostic state-selected seed-bias tables retain AUC 0.74–0.87 across the two families.

The deployed watermark does not survive this replacement. The diagnostic changes noise invariance and key-space design and is not the deployed method. e-distillatione-distillation-setup

6.3 Ablations and diagnostic examples

Read component removals and qualitative examples within their stated evaluation conditions.

Table 2. Unknown hidden-channel completion is the weakness of reverse ODE recovery. Original paper, p. 7 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Begin with the full-observation rows, where both methods receive a complete raw-action endpoint. Then move to partial observation, where the π0.5/LIBERO verifier sees seven of thirty-two channels. Reverse ODE must decide how to fill the other channels: no completion, correct zero padding, an incorrect unit value, or a secret random pad. MAP instead scores only the visible projection. Compare the AUC column within the partial block rather than mixing it with the full block. Cohen's d measures separation, and the last column reports recovered latent norm where available. A dash means that norm was not reported, not that it was zero. e-recoverye-map

What it supports. Partial MAP reaches AUC 0.775 without knowing the padding convention. Correctly matched ODE reaches 0.969, but incorrect and secret padding reduce it to 0.678 and 0.508. The useful result is robustness to a hidden implementation convention, rather than superiority over an oracle with the correct completion.

Where the evidence stops. These are per-chunk separation measurements on fifty marked and fifty plain rollouts of one clean policy. They are not sixteen-rollout decision AUCs, and the table does not show uncertainty or prove unique latent recovery.

Table 7. Global alignment restores the fingerprint after short fixed delays. Original paper, p. 17 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Read across a delay row before moving down to longer delays. The left half fixes lag at zero; the right half enables one global synchronization search. Each half starts with single-episode AUC, followed by TPR at group budgets sixteen, thirty-two and sixty-four. These columns are different metrics and should not be interpreted as one continuous AUC curve. The table labels the estimated lag τ*, whereas Equation (8) uses ℓ*. The equation chooses a lag from the owner-key response over the group, then scores both owner and decoy keys at that common alignment. e-synce-scoree-attack-table

What it supports. For a one-step delay at sixteen episodes, TPR rises from 0.28 to 1.00. A three-step delay remains harder: search gives 0.53 at sixteen episodes and 1.00 at sixty-four, whereas lag-zero remains at 0.00. This isolates a timing correction in the reported Table 7 comparison.

Where the evidence stops. Table 6 also says delay uses synchronization, yet gives two-/three-step TPR of 0.41/0.26 at sixteen episodes, versus 0.96/0.53 here. No protocol difference explains this conflict. The crop is faithful; the retained numbers are explicitly Table 7 values.

Figure 9. A copied behavior can survive while the deployed provenance mark disappears. Original paper, p. 13 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Read the left panel for π0.5 and the right for LingBot-VA. The horizontal coordinate counts entries in the diagnostic seed-bias lookup table. Blue circles, labeled AUC (DC), measure cross-student detection for that biased seed-control diagnostic; orange triangles report its student success rate. These are distinct quantities sharing a vertical scale. Red crosses show the deployed zero-mean fingerprint after behavior cloning, close to the dashed chance line. Appendix .2 explains the distinction: the diagnostic chooses a persistent bias from the current state, while students are evaluated without the owner's keyed sampler. The red crosses do not define the deployed method's key-space size. e-distillatione-distillation-setupe-threat

What it supports. The deployed mark gives post-cloning AUC 0.504 and 0.508, effectively chance in these tests. The diagnostic retains AUC 0.74–0.87 across both panels by giving the student a repeatable pattern to learn. For LingBot, the diagnostic SR labels are 0.90, 0.92 and 0.95.

Where the evidence stops. The diagnostic sacrifices the deployed design's noise-law invariance and key-space properties. Its improved inheritance is not evidence that the original watermark resists distillation. Orange utility values describe diagnostic students, not the deployed-mark students.

7. Analysis & limitations

7.1 What the evidence leaves open

Reader analysis

The binary threshold calibrated at nominal 1% yields plain-rollout FPR 2.0% with one episode and 2.4% with sixteen for LingBot/RoboTwin. Good AUC cannot establish correct fixed-threshold false-attribution control. e-calibration

Reader analysis

Adaptive LoRA stress testing leaves group detection strong, but the LingBot/RoboTwin-10 identification example falls from DIR 0.999 at 82% task success to 0.570 at 66%, while group AUC remains 0.999. This uses a different utility anchor from Table 1 and must remain a separate setting. e-adaptivee-main

Reader analysis

Security tails of 10^{-6} and 10^{-9} are fitted extrapolations, not measured guarantees. The empirical 10^{-3} operating point comes from a finite rollout pool and bootstrapped groups; 128-bit keys do not imply 128-bit statistical false-attribution resistance. e-security

Reader analysis

The source does not reconcile its i.i.d. Gaussian reference construction with the experimental description of four Gaussian tones in 1–2 Hz. Figure 2 labels output β=0.3 while Section 5.2 specifies 0.05. Tables 6–7 disagree on delayed detection, and Figure 10's EMA 0.2 curve rises near 0.9 despite the text's below-0.2 boundary. These are unresolved source inconsistencies. e-injectione-protocole-spectrume-attack-tablee-synce-budget-conflict

7.2 Questions for discussion

  1. Does independent no-injection calibration control false attribution after task or controller changes?
  2. How exactly are the stated band-limited tones reconciled with i.i.d. Gaussian reference coordinates?
  3. What audit cost and useful threat coverage remain once an adversary can train a replacement sampler?

8. Reproducibility audit

8.1 Requirements and known gaps

Reader analysis

Required artifacts include the exact policy checkpoints and samplers, channel projection, observation/instruction records, episode nonces, chunk timing and post-processing model. Appendix .1 supplies key hyperparameters but omits MAP step count, learning rate, restart count and compute. The supplied paper does not specify a complete software/hardware configuration, a uniform rollout-pool accounting, a clear held-out calibration split or a physical-robot deployment protocol. e-mape-protocole-hyperparameterse-task-aggregation

Reader analysis

Reader-proposed priorities are to repeat recovery with matched and mismatched hidden padding, and independently test decoy thresholds on held-out plain rollouts. Resolve the reference-generation and delay-table inconsistencies before treating either as an exact replication target. The edition specifies controls and falsifiable outcomes. e-recoverye-calibratione-synce-injectione-protocol

8.2 Proposed reproduction checks

The following checks are proposals motivated by the paper. They have not been run as part of this reading.

Check 1: Recover the key while hiding the training padding convention

Reader-proposed check, not performed: use one fixed π0.5/LIBERO checkpoint and the same marked/plain rollout pairs for all arms. Keep visible channels, nonces, key bank and recovery compute fixed. Compare MAP with reverse ODE using correct zero padding, incorrect constant padding and undisclosed random padding; retain a full-observation control. Report held-out per-chunk AUC, recovered norms, failure rates and wall-clock cost, then evaluate sixteen-episode decisions separately. The mechanism predicts that ODE's advantage depends on correct completion while MAP changes less when only the hidden completion is altered. A comparable MAP collapse, or stable ODE under wrong completion, would weaken that explanation. First specify the paper's omitted optimizer/restart settings and resolve the reference-construction ambiguity. e-recoverye-mape-hyperparameterse-injectione-protocol

Check 2: Test whether decoys calibrate unseen unmarked services

Reader-proposed check, not performed: split plain and marked LingBot/RoboTwin rollouts by episode into calibration and held-out evaluation pools, with disjoint nonces. Compare the paper's decoy-calibrated binary threshold against a threshold calibrated on plain rollouts; keep the model, recovery, key bank and group budgets one and sixteen fixed. Repeat clean, jitter and fixed-delay conditions, with and without global lag search. Freeze thresholds before evaluating held-out plain FPR and marked TPR, and report uncertainty from independent episode groups. The falsifiable target is nominal one-percent error control without losing the carried-key signal. Persistent excess FPR would show that score normalization or alignment selection needs additional calibration, even when AUC remains high. e-calibratione-scoree-protocole-sync

8.3 Reading coverage

Visual audit: The title page, all method equations, Figures 1–13, Tables 1–7 and Appendix .1–.8 were visually inspected alongside the complete supplied text. All six final crops were inspected for readability and complete labels. The audit includes uncropped pages supporting hyperparameters, stress-test training, security analysis and proposed checks. Figure 1's plus symbol is interpreted through Eq. (1); Table 7's lag notation and its conflict with Table 6 are disclosed. The report also records the Figure 2 amplitude mismatch, the Figure 10 smoothing inconsistency and the unresolved reference-law description. No separate supplement, code inspection or experimental reproduction is included.

PDF pages inspected for this edition: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19. Appendix coverage: reviewed.

Original figures and tables remain the work of the source’s authors. Extractions preserve their scientific content; any HTML wrapper layout is disclosed with each figure. The surrounding reading notes are our own.

Text reading scope & known omissions
  • Title and abstract
  • 1 Introduction
  • 2 Related Work
  • 3 Threat Model
  • 4 Method; 4.1 Fingerprint Injection; 4.2 Fingerprint Verification
  • 5 Verification Evaluation; 5.1–5.6
  • 6 Identification Evaluation
  • 7 Security Analysis
  • 8 Discussion
  • 9 Conclusion
  • Ethics Considerations; Generative AI Usage Considerations
  • References
  • Appendix .1 Method Hyperparameters
  • .2 Distillation Stress-Test Details
  • .3 Adaptive-Removal Stress-Test Details
  • .4 Budget Sweeps for Weak Cells
  • .5 Synchronization-Search Ablation
  • .6 Same-Task vs. Cross-Task Aggregation
  • .7 Detailed Clean Identification Curves
  • .8 Detailed Identification Robustness

Outside the original text pass

  • Text extraction does not reconstruct figure images; inspect the retained PDF for figures and equation/table layout.
  • Separate supplemental material availability has not been fully verified.
  • The reviewed artifact is arXiv:2606.23574v1, dated 22 June 2026. Its title and four authors match the catalog; no other revision was supplied or compared.
  • The acquisition manifest notes that text extraction does not reconstruct figure images; the retained PDF was therefore visually inspected on all 19 pages.
  • Separate supplemental material availability has not been fully verified; no separate supplement was supplied.
  • The linked code was not inspected, and no experiments were reproduced.

The visual audit above records the subsequent illustrated pass.

8.4 Traceable evidence

e-identityPDF p. 1, title block and arXiv marginInspect

Title matches the supplied observed title. Authors are Yule Liu, Shuai Liu, Jiaheng Wei and Xinlei He; the margin identifies arXiv:2606.23574v1, 22 Jun 2026. The three listed institutions match the title metadata.

Go to primary source ↓
e-architecturePDF p. 2, Figure 1; p. 3, contributions and Section 2Inspect

Figure 1 links keyed sampling, policy generation, partial actions, MAP and key aggregation. Section 2 writes a direct VLA generator and a WAM future-scene/action composition. The contribution requires no watermark weight retraining.

Go to primary source ↓
e-threatPDF pp. 3–4, Section 3, owner capability and scope statementInspect

Default adversaries can invoke and post-process a sealed service. Audit uses executed commands and the owner's base model; the protected keyed sampling process must remain active. A fresh student with a new sampler is outside this default scope.

Go to primary source ↓
e-injectionPDF p. 4, Section 4 notation and Section 4.1, Eq. (1); p. 12, Eq. (12)Inspect

Key, public nonce and chunk determine references and selected positions. At most m chunks are selected with maximum gap P. The Gaussian-law argument assumes independent standard-Gaussian base and reference vectors; Eq. (12) describes coordinate-wise BLAKE2B-to-Gaussian expansion.

Go to primary source ↓
e-mapPDF p. 5, Section 4.2, Eqs. (2–3) and MAP recovery paragraphInspect

The observation model projects onto executed channels and includes post-processing. MAP balances observation fit with a squared latent norm, using multiple random starts and the smallest final objective. The paragraph gives no iteration, learning-rate or restart count.

Go to primary source ↓
e-scorePDF pp. 5–6, Eqs. (4–8), synchronization search and Proposition 1Inspect

A flattened inner product supplies the score; decoys estimate episode mean/scale, and standardized scores sum over rollouts. One global lag maximizes owner-key response and is reused for decoys. The rate law assumes independent episodes and Gaussian approximation.

Go to primary source ↓
e-protocolPDF pp. 6–7, Section 5.1, evaluation design and common protocolInspect

Executed/raw counts are 7/32, 7/30, 14/32 and 16/30 for the four named cells. The main protocol uses up to five chunks, β=1, four Gaussian tones in 1–2 Hz, and 32 decoys. Groups and confidence bands are formed by bootstrapping episode pools.

Go to primary source ↓
e-mainPDF p. 7, Table 1, all four rows and captionInspect

All clean group AUC/TPR values are 1.000 at budget 16. The table provides the retained single-episode AUCs and plain/marked SR pairs; canonical-attack averages include π0.5/LIBERO TPR 0.840. Utility entries have no uncertainty intervals.

Go to primary source ↓
e-recoveryPDF p. 7, Section 5.3 and Table 2; p. 8, continuationInspect

Same 50 marked/50 plain π0.5/LIBERO rollouts compare recovery rules. Table 2 gives full MAP/ODE AUC 0.994/0.809 and partial MAP 0.775 versus ODE values 0.492, 0.969, 0.678 and 0.508 across completion assumptions.

Go to primary source ↓
e-calibrationPDF p. 8, Figure 4; p. 9, Section 5.4 continuationInspect

At nominal 1%, plain FPR for budgets 1/16 is LingBot/LIBERO 0.8%/0.3%, LingBot/RoboTwin 2.0%/2.4%, π0.5/LIBERO 0.0%/0.0%, and π0.5/RoboTwin 0.3%/0.1%. The text acknowledges anti-conservative calibration in LingBot/RoboTwin.

Go to primary source ↓
e-spectrumPDF p. 6, Figure 2 legend; p. 7, Section 5.2Inspect

Figure 2's output-sine legend states β=0.3, whereas the output baseline setup states β=0.05. The spectral and notch-sweep plots support localization of that output mark, but the displayed amplitude and setup value differ.

Go to primary source ↓
e-variantsPDF p. 9, Section 5.5 and Table 3; p. 10, compression continuation; p. 11, Table 5Inspect

LoRA descendants give group detection TPR 1.00. π0.5/RoboTwin action-side prune30/int8 have TPR 0.56/0.94 at budget 16. Weight edits are owner variants; base recovery is retained. Identification on pruning is weaker than binary detection.

Go to primary source ↓
e-identificationPDF p. 10, Section 6 decision rules; p. 11, Table 4; p. 18, Figure 12Inspect

The gallery has one true key and 32 decoys. Open-set rejection uses maxima on plain impostors. Table 4 gives all rank-1 and DIR values retained here, and Figure 12 displays their budget dependence.

Go to primary source ↓
e-securityPDF p. 12, Section 7, Eqs. (9–12) and Figure 7Inspect

The study uses 1,024 decoys, 300 fingerprinted episodes and 200,000 bootstrap groups of size 16. The 10^{-3} false-key point is empirical; 10^{-6}/10^{-9} points use fitted tails. Key entropy is 128 bits, distinct from statistical tail risk.

Go to primary source ↓
e-adaptivePDF p. 13, Eq. (13), Figure 8 and adaptive-removal text; p. 16, Appendix .3Inspect

Adaptive LoRA minimizes recovered-key response while penalizing task/action drift and parameter changes. LingBot/RoboTwin-10 has the stated 82%/66% utility and 0.999/0.570 DIR pair, with group AUC 0.999 after the borderline attack.

Go to primary source ↓
e-distillationPDF p. 13, Figure 9 and Section 8Inspect

Deployed-mark cross-student AUC falls to 0.504/0.508. Small state-selected seed-bias tables yield AUC 0.74–0.87. The diagnostic trades Gaussian-law invariance and large key space for a persistent learnable pattern; orange curves report diagnostic student SR.

Go to primary source ↓
e-hyperparametersPDF p. 15, Appendix .1Inspect

β=1, m=5; P=1 for π0.5, 6 for LingBot/LIBERO and 2 for LingBot/RoboTwin. λ_z=1; σ_obs=10^{-4} for π0.5 and 10^{-3} for LingBot. No optimizer iteration/restart count or hardware configuration is provided here.

Go to primary source ↓
e-distillation-setupPDF p. 15, Appendix .2Inspect

π0.5 student training relabels a fixed-observation demonstration corpus and trains attention-only LoRA for 1,500 steps. Diagnostic lookup sizes are 20, 80 and 160; state is quantized and hashed to select a bias. LingBot uses LIBERO-10, ten tasks and ten episodes per task.

Go to primary source ↓
e-syncPDF p. 16, Appendix .5; p. 17, Table 7 and captionInspect

Table 7 compares lag-zero with global τ* search. At budget 16, delay 1/2/3 TPR changes from 0.28/0.02/0.00 to 1.00/0.96/0.53; at 64, search yields 1.00 in all rows. Eq. (8) calls the lag ℓ*.

Go to primary source ↓
e-attack-tablePDF p. 17, Table 6, π0.5/LIBERO and π0.5/RoboTwin columns and caption; p. 9, output-level resultsInspect

Table 6 reports π0.5/LIBERO EMA-lo and jitter-hi TPR 0.09/0.17 at budget 64; delay-2/3 TPR is 0.41/0.26 at 16 despite the search-enabled caption. These differ from Table 7. Compression is restricted to action-side parameters for π0.5/RoboTwin.

Go to primary source ↓
e-budget-conflictPDF p. 16, Figure 10; p. 9, Section 5.5 boundary statement; p. 17, Table 6Inspect

Figure 10 labels an EMA 0.2 curve ending near 0.9 at budget 64, whereas Section 5.5 says the strongest smoothing/jitter stay below 0.2 and Table 6 reports EMA-lo 0.09. Strength mapping and this discrepancy are not resolved by the supplied paper.

Go to primary source ↓
e-task-aggregationPDF p. 16, Appendix .6; p. 17, Figure 11 captionInspect

Same-task and cross-task aggregation nearly overlap on LingBot. The caption says π0.5/LIBERO has one episode per task and π0.5/RoboTwin one evaluation task, so their same-task panels are not meaningful comparisons.

Go to primary source ↓

8.5 Primary sources

Scroll across the image to inspect details. Press Esc to close.