Paired Exact-Reset Evaluation of a Prediction-Derived Medium-to-Full World-Model Cascade
1. Paper overview
In one sentence: A Medium-derived prediction interface can improve selective physical decisions enough to pay for sequential overhead, while still taking longer than fixed prediction policies. e-probleme-pairinge-controle-v107e-operatinge-price
| At a glance | What to know |
|---|---|
| Research problem | Source description Higher predictive capacity can change the ordering of task-relevant actions unfavorably. The evaluation question is whether information available after Medium predicts when Full's physical decision improvement justifies a second complete computation. Average prediction accuracy or latent distance alone cannot settle that question. e-probleme-recoding |
| Core mechanism | Source description Candidate-complete exact resets make Medium-minus-Full physical regret observable on the same query, separating model selection from changes in sampled states or available actions. e-pairing |
| A key reported result | Prospective PushT V107 prediction versus current-DINO/action routing: −0.002549; state-clustered 95% interval [−0.002867, −0.002238]; one-sided 95% upper bound −0.002286. All three checkpoint-pair point effects are negative. Prediction-minus-control priced physical decision cost; lower is better.. 1,600 fresh states; 39 tasks; five actions; horizon 30; three equally weighted frozen checkpoint pairs; lambda=0.002. Regret 0.02849 versus 0.03110; latency 0.664 versus 0.630 ms; escalation 47.2% versus 66.2%. The prespecified gate passes. The reported 7.9% priced-cost reduction comes from lower physical regret despite slower execution, and supports the tested interface comparison. e-controle-pushte-v107e-v107-tablee-estimands |
| Reading caution | Source description All outcomes concern controlled simulator action sets and task losses. No closed-loop robotics deployment is evaluated; related DINO-style predictors and three fixed checkpoint pairs do not establish generality across world-model families or training randomness. e-limits |
Core contributions
- Source description
Candidate-complete exact resets make Medium-minus-Full physical regret observable on the same query, separating model selection from changes in sampled states or available actions. e-pairing
- Source description
Frozen V106 and PyBullet audits are extended by V107, a separately sealed PushT confirmation against a stronger current-state/action control. The paper reports an uncertainty repair and explicit archival boundaries. e-controle-inferencee-archive
Figure 1. Exact resets supply offline supervision; deployed routing uses information exposed by Medium. Original paper, p. 2 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. Follow the upper arrows from query Q, containing observation, task and candidates, through Medium's predicted consequences to the router. The right box retains action A_M when pi=0, or evaluates Full and takes A_F when pi=1. The lower branch executes all candidates from exact resets, constructing paired outcomes and a benefit label; its dashed upward arrow is explicitly for fitting and auditing. The red boundary describes information unavailable to the router before escalation. It does not mean that Full is never evaluated: Full is the conditional second-stage computation shown in the rightmost box. e-diagrame-pairinge-ledgere-routere-restrictions
What it supports. The separation makes the central test possible: actual outcomes determine whether switching helped, while the deployed selector must infer useful switching decisions from its restricted interface. This is action selection from a common candidate set, with physical outcomes evaluated in simulators, rather than a learned joint action-and-future output.
Where the evidence stops. The figure labels B as net benefit, but §3.1 specifies fitting gross regret difference and selecting a validation threshold. Read B as the cost-ledger quantity; the executed score is not asserted to be a calibrated Bayes net-benefit estimate.
2. Motivation
2.1 The problem and the proposed response
Higher predictive capacity can change the ordering of task-relevant actions unfavorably. The evaluation question is whether information available after Medium predicts when Full's physical decision improvement justifies a second complete computation. Average prediction accuracy or latent distance alone cannot settle that question. e-probleme-recoding
2.2 What this reading follows
A larger world model can choose a worse action even when it is the more capable predictor on average. This paper makes that possibility measurable: execute every candidate from the same simulator state, score both predictors' chosen actions against those outcomes, and learn when switching helps. The strongest test compares prediction routing with a current-DINO/action router on a second, prospectively sealed PushT bank. Read the figures as an argument about information, selection and accounting. The positive result is a reduction in latency-priced physical regret; the operating tables and price sweep explain why that result does not establish faster control or a general deployment frontier. e-probleme-pairinge-controle-v107e-operatinge-price
3. Research context
We place the paper in the collection through its world–action interface. The catalog labels and the reading’s assessment are shown separately.
| Catalog dimension | Recorded classification |
|---|---|
| Major category | Evaluation metrics |
| Architecture | Not applicable |
| Prediction paradigm | Not applicable |
| Quadrant | Not applicable |
3.1 Evidence-based assessment
Supports the recorded classification
The recorded evaluation-metrics/protocol classification is supported: the central contribution is a paired decision-loss estimand and frozen statistical audit. The experimental cascade combines separately frozen predictors with a router and task scorer; it supplies no unified joint future/action predictor or inverse-dynamics action extraction. Architecture, prediction paradigm and quadrant are appropriately not applicable to this protocol entry. e-probleme-diagrame-pairinge-router
This is the collection’s architectural analysis, not a new related-work survey. Benchmark comparisons and their protocols appear in Section 6.
4. Problem formulation
4.1 Inputs and outputs
| Inputs | Outputs |
|---|---|
|
|
4.2 Equations and their role
5. Method in detail
5.1 Turn model disagreement into a paired physical decision target
Begin with one observed simulator state, a downstream task and a finite set of candidate actions. Medium and Full each predict consequences for those same candidates, and the same task loss turns each predictor's outputs into an action ranking. Now execute every candidate from the identical reset state. The resulting physical outcomes provide both models' realized action losses and the best achievable loss within that candidate set. Subtracting that baseline gives two directly paired regrets. Their difference can supervise a router without conflating predictor quality with different state samples. This construction also defines the boundary: the oracle is best among the tested candidates, and its outcomes are offline supervision. It is not an unrestricted action oracle or information available when the deployed router chooses whether to escalate. e-pairinge-routere-restrictions
5.2 Price the complete cascade before interpreting its benefit
A prediction-interface router first pays for Medium and for exposing the routing features. Escalation then adds a complete Full prediction and scoring pass. The net-benefit quantity B subtracts that additional priced latency from Medium-minus-Full physical regret; its conditional expectation defines the ideal routing decision. The implemented regressors fit gross regret differences, however, and validation chooses strict thresholds rather than assuming perfect calibration. Comparing with standalone Full requires another accounting step: Full alone avoids the initial Medium computation. The paper's signed latency-gap term records that difference. Consequently, a router can escalate less often and still take longer. Table 4 supplies a concrete example: the prediction interface's physical-regret gain offsets its latency disadvantage at the frozen price, yielding a better objective without a faster policy. e-ledgere-routere-theorye-v107-table
5.3 Separate interface evidence from repeated use of the same states
V106's task-only control almost always chooses Full, so its defeat leaves a plausible current-state explanation open. V107 addresses that explanation with a dimension-matched current-DINO/action router and a fresh bank generated after the comparison is sealed. The control is deliberately favored in the latency ledger. For inference, multiple tasks and checkpoint effects attached to one physical state are averaged before resampling states; they do not become independent observations merely because they occupy separate rows. This distinction also explains the PyBullet repair: its original cellwise bootstrap broke shared-state pairing across checkpoints. Repaired intervals preserve that pairing and retain the original point estimates. The resulting claims remain conditional on the fixed tasks, banks and checkpoint pairs, with wider checkpoint resampling displays treated as sensitivity rather than population coverage. e-v106e-controle-inferencee-estimands
5.4 Training and inference
During training
PushT predicts over a 64-D PCA of pooled DINOv2 ViT-S/14 features. Medium uses two 256-wide GELU layers, Full three 512-wide layers. AdamW training uses 4,950 action-conditioned rows, selection 900 rows, then capacity-specific ridge physical calibration uses 180 states. Router training and validation use 5,200 and 1,200 states, each crossed with 39 tasks. e-pusht
PyBullet uses residual token Transformers over pooled DINOv2 features: Medium width/heads/layers 192/4/2 and Full 384/6/3. AdamW training runs 100 epochs on 5,000 states, followed by physical calibration on 1,000 states. The 59-dimensional task–prediction–regime interface was chosen before confirmation; Table 3 leaves PyBullet parameter and MAC counts unspecified. e-pybullet
World predictors and calibrators remain frozen during routing evaluation. V107's scaler/PCA uses V106 training states, and thresholds use V106 validation states; its two fresh 800-state shards provide no fitting or selection feedback. e-controle-pusht
During inference
The prediction router always pays for Medium first. A strict score-above-threshold decision either keeps its action or pays for Full's complete prediction and scoring pass. The input-routing controls can bypass Medium; V107 additionally charges no DINO encoder latency. There is no evaluated closed-loop feedback policy. e-routere-controle-limits
5.5 Implementation flow
- Execute paired candidate outcomes offline
Each predictor scores the same actions through the task loss. PushT resets to its recorded seeded base state before each rollout; PyBullet restores a saved simulator state. Every candidate outcome supplies the physical best-in-set baseline. Duplicate replay checks sample three states per PushT shard and one candidate per checked state, so determinism validation is not exhaustive. e-pairing
- Restrict the routing interface
Medium's calibrated consequence and the task feed a router; true futures, realized regret, oracle actions, Full predictions and explicit simulator state are excluded. V107 compares 22 task features plus 43 prediction features against the same tasks plus a 43-dimensional projection of current 64-D DINO and all five actions. Both inputs have 65 dimensions. e-controle-restrictions
- Separate fit, selection and confirmation
Routers regress gross paired regret improvement using 128/64-unit ReLU MLPs. Strict thresholds are selected on validation from constant endpoints and 101 empirical score quantiles. Figure 1 depicts the net-benefit ledger, whereas §3.1 specifies gross-target fitting. V107 freezes routers, projection, thresholds and its primary comparison before generating fresh outcomes. e-diagrame-routere-control
6. Experiments & results
This paper evaluates when a Medium predictor's imagined consequences help select between separately frozen Medium and Full computations. Exact simulator resets supply paired physical decision losses for identical candidate actions. A prospective PushT confirmation finds lower latency-priced regret than a dimension-matched current-DINO/action router, although prediction routing is slower. The contribution is an evaluation protocol with explicit cost accounting and clustered inference; its evidence concerns controlled candidate-set decisions, without establishing closed-loop robotics value.
The paper refers to a development interface ablation but supplies no standalone numerical component-removal table for it. This edition uses the original capacity-reversal and compute-price diagnostics as its ablation-section visuals. Figure 1 is the protocol's information-flow schematic; the paper's contribution is evaluation rather than a new neural architecture. No closed-loop experiment is reported. e-probleme-pybullete-restrictionse-limits
6.1 Read the original evidence
Figure 3. A fresh PushT bank favors the tested prediction interface for every fixed checkpoint pair. Original paper, p. 9 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. Start at the dashed zero line: points to its left favor prediction routing. The aggregate row averages 39 tasks and the three fixed checkpoint-pair effects within each of 1,600 fresh states, then uses state-clustered bootstrap inference. The other rows expose the three checkpoint-pair effects separately, making it possible to check that the aggregate is not hiding a positive point effect. The control has the same total input dimension and learner family, receives current DINO plus all candidate actions, routes before Medium and pays no DINO encoder latency. The comparison and its pass conditions were frozen before these outcomes were generated. e-controle-v107e-estimandse-limits
What it supports. The aggregate effect is −0.002549 at lambda=0.002, with a 95% interval of [−0.002867, −0.002238]. The one-sided upper bound is −0.002286, and each fixed checkpoint-pair point is negative. Together these satisfy the prespecified V107 gate and support incremental routing information in this tested prediction interface.
Where the evidence stops. The bars condition on fixed tasks and checkpoint pairs; they are not uncertainty over a population of trained models. Beating this dimension-matched control does not establish causal sufficiency of predicted futures or superiority to every current-state representation.
Table 4. Lower physical regret outweighs a small latency disadvantage in V107. Original paper, p. 9 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. Read across each policy row before comparing the last column. Regret measures the physical loss of the selected action relative to the best candidate; decision cost adds latency priced at lambda=0.002. The bottom row subtracts the current-DINO/action control from prediction routing. Its negative regret difference and positive latency difference therefore pull the objective in opposite directions. Escalation falls from 66.2% to 47.2%, a difference of 19.0 percentage points, yet latency rises. The explanation is sequential accounting: prediction routing always runs Medium, whereas the input control can select Full without paying for Medium first. e-v107-tablee-v107e-controle-ledgere-pairing
What it supports. The table reports regret 0.02849 versus 0.03110 and latency 0.664 versus 0.630 ms. The text's more precise 0.002617 regret improvement exceeds the 0.000068 priced latency disadvantage, yielding the −0.002549 primary cost effect. The reported relative decision-cost reduction is 7.9%, despite slower raw execution.
Where the evidence stops. Displayed means and differences are rounded, so their last digits need not reproduce the exact effect. A lower escalation percentage is not itself compute saving when the compared policies pay different first-stage costs.
Table 5. The original routers improve the priced objective while taking longer than either fixed policy. Original paper, p. 11 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. Compare policies within each environment's three-row block, using the latency and regret columns together. The PushT rows refer to V106's original operating point, not the later V107 comparison with current DINO. All decision costs use lambda=0.002, while the environment and checkpoint weights follow their frozen estimands. PyBullet's router escalates only 17.8% of the time, but its always-paid interface costs leave it slower than Fixed Full. The same pattern appears in PushT at 42.0% escalation. The table therefore provides a direct check on whether a selective policy is buying lower physical regret, lower latency, or both. e-operatinge-timinge-ledgere-estimandse-limits
What it supports. PushT router latency is 0.651 ms versus 0.281 for Medium and 0.364 for Full; its decision cost is nevertheless lower at 0.03037. PyBullet similarly reaches cost 0.02573 while taking 1.320 ms versus 0.909 and 1.103 ms. Both reported operating points exchange additional time for improved physical choices.
Where the evidence stops. Timing combines measured GPU model forwards and CPU components, rather than synchronized end-to-end deployment. These means do not establish robot success, portable latency, or closeness to the Bayes interface oracle.
6.2 Results and evaluation conditions
| Task & protocol | Reported result | Comparison & interpretation |
|---|---|---|
| Prospective PushT V107 prediction versus current-DINO/action routing 1,600 fresh states; 39 tasks; five actions; horizon 30; three equally weighted frozen checkpoint pairs; lambda=0.002. | −0.002549; state-clustered 95% interval [−0.002867, −0.002238]; one-sided 95% upper bound −0.002286. All three checkpoint-pair point effects are negative. Prediction-minus-control priced physical decision cost; lower is better. | Regret 0.02849 versus 0.03110; latency 0.664 versus 0.630 ms; escalation 47.2% versus 66.2%. The prespecified gate passes. The reported 7.9% priced-cost reduction comes from lower physical regret despite slower execution, and supports the tested interface comparison. e-controle-pushte-v107e-v107-tablee-estimands |
| Initial PushT V106 frozen routing audit 1,600 sealed states, 39 tasks, three fixed checkpoint pairs; lambda=0.002; five primary references. | Versus Medium: −0.004308 [−0.004759, −0.003857]; versus Full: −0.002577 [−0.003001, −0.002159]; intervals are state-clustered 95%. Router-minus-reference priced decision cost. | Versus task-only: −0.002620 [−0.003037, −0.002213]. The task-only router escalates on 99.1453% of rows under equal checkpoint weighting. The weak task-only selector motivated V107. V106's matched-count random benchmark is descriptive, not a hypothesis-test p-value. e-pushte-v106e-primary-comparisons |
| Controlled-PyBullet V104E confirmation Six equally weighted 2,000-state banks; 81 tasks; three fixed checkpoint pairs; lambda=0.002. | Versus Medium: −0.003677 [−0.003805, −0.003553]; versus Full: −0.003746 [−0.003880, −0.003613]; repaired state-clustered 95% intervals. Router-minus-reference priced decision cost. | Matched-rate margin: −0.002800 [−0.002902, −0.002698]; matched random: −0.003668 [−0.003773, −0.003565]. All four adjusted primary upper bounds remain negative after the bootstrap repair. This tests a composite task–prediction–regime interface, without isolating prediction from regime. e-pybullete-inferencee-primary-comparisonse-estimands |
| PushT compute-price sensitivity V106's declared price grid; thresholds selected separately on validation for each price. | Against Full: −0.000385 at lambda=0.01; approximately +0.0030 at 0.025 and +0.0074 at 0.05. Router-minus-reference decision cost. | All five listed references beat the router at 0.025 and 0.05. V107's separate fixed-threshold comparison remains negative against its DINO/action control across the declared grid. A low-price operating advantage is established. The two sensitivity protocols cannot be merged into a continuous deployment frontier. e-pricee-v107-price |
6.3 Ablations and diagnostic examples
Read component removals and qualitative examples within their stated evaluation conditions.
Figure 2. Both capacities win on substantial subsets of the finite audit. Original paper, p. 8 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. Read each horizontal bar as a partition of physical states, with blue indicating lower decision cost for Medium and red for Full. The caption defines the aggregation: downstream tasks and three frozen checkpoint effects are averaged within each unique state before assigning its winning capacity. Thus 45.1% for PushT describes a fraction of states under that averaging rule, not the probability that Medium wins an individual action, task or new training run. The panels describe the original finite PushT and controlled-PyBullet audits. They motivate selective computation by showing that choosing Full everywhere can discard useful Medium decisions. e-reversalse-restrictionse-limits
What it supports. Medium wins on 45.1% of PushT states and 51.9% of PyBullet states; Full wins on their complements. Realized state-level heterogeneity creates an opportunity for a selector that could identify the favorable regions. Whether the allowed prediction interface actually exposes that opportunity must be tested with frozen routers.
Where the evidence stops. This is a diagnostic of realized capacity reversals, not a component-removal ablation or proof of Blackwell incomparability. The router never receives the realized winning-capacity labels at deployment.
Figure 6. The original PushT advantage disappears when computation receives a sufficiently high declared price. Original paper, p. 12 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. Use the horizontal axis as the price assigned to measured milliseconds and the dashed horizontal line as equal decision cost. Each colored curve compares the candidate router with a different reference listed in the legend. Negative points favor the router. Read the marked prices individually: the protocol selects thresholds on validation separately for each price, so the connected points do not represent one unchanged policy as price varies continuously. At the tested low prices the gain in physical regret exceeds the priced overhead; at the two highest prices the balance reverses against all five references. e-pricee-v107-pricee-limits
What it supports. The candidate is favorable at the tested prices through 0.01, although its Full comparison is already narrow there: −0.000385 in the text. At 0.025 and 0.05 every displayed difference is positive. This locates the supported operating region and prevents interpreting the primary low-price gain as universal efficiency.
Where the evidence stops. V107's Appendix E sensitivity holds thresholds fixed and compares against current-DINO/action routing; it is a different protocol. Neither display establishes a continuous Pareto frontier, and superiority to another router alone does not imply superiority to fixed policies.
7. Analysis & limitations
7.1 What the evidence leaves open
All outcomes concern controlled simulator action sets and task losses. No closed-loop robotics deployment is evaluated; related DINO-style predictors and three fixed checkpoint pairs do not establish generality across world-model families or training randomness. e-limits
The router is slower than both fixed policies in both original audits. The realized-loss clairvoyant lower bound uses unavailable outcomes and zero router overhead, so its gap is not Bayes-interface regret. Empirical capacity reversals likewise do not establish Blackwell incomparability. e-operatinge-reversalse-limits
Population uncertainty assumes physical-state cluster sampling and fixed tasks, banks and checkpoints. PyBullet's original cellwise bootstrap broke crossed-state pairing; the repair changes uncertainty only. Crossed-checkpoint sensitivity quantiles do not provide population coverage over checkpoint training. e-inferencee-estimands
7.2 Questions for discussion
- Would the prediction-interface advantage persist with a stronger current-state representation and separately synchronized end-to-end timing?
- How would query-dependent incremental latency change the benefit target and validation-selected threshold policy?
8. Reproducibility audit
8.1 Requirements and known gaps
The paper describes released policy-facing arrays, router checkpoints, thresholds, manifests and frozen code sufficient to recompute reported statistics. Raw observations, all-candidate outcomes, DINO banks, upstream predictor checkpoints and the original HPC environment are absent from that archive. Self-authored seals do not externally certify chronology. e-archive
Replay requires exact serialized thresholds cast to NumPy float32, strict score>threshold routing and first-minimum argmin ties. A different float parser plus float64 promotion changed 6,400 high-price decisions. Preserve shared-state pairing in all bootstrap resamples. e-replaye-archivee-estimands
Latency uses H100 world-model forwards and single-threaded CPU task/router components; the CPU SKU is unrecorded. PushT uses 200 warmups, seven rounds of 1,000 repeats and 512 sampled queries. Synchronized deployment timing and several training details, including learning rates and explicit predictor-loss formulas, are not specified in this PDF. e-timinge-pushte-pybullete-limits
8.2 Proposed reproduction checks
The following checks are proposals motivated by the paper. They have not been run as part of this reading.
Check 1: Replay strict-threshold routing and its cost ledger
Reader-proposed check, not performed: using the described released scores, threshold strings and latency components, replay V106 with thresholds cast directly to NumPy float32 and strict score>threshold decisions. Compare masks and costs with archived outputs across the declared price grid. As a controlled perturbation, parse the thresholds through the alternative default float parser and promote scores to float64 while holding all data and cost weights fixed. The reference replay should recover the archived masks and costs; the perturbed replay should localize the documented high-price decision changes. If the reference masks disagree, resolve score serialization and action-tie semantics before drawing scientific conclusions from the cost sweep. e-replaye-archivee-pricee-ledger
Check 2: Recompute uncertainty while preserving crossed physical states
Reader-proposed check, not performed: use the described policy-facing arrays to reconstruct V107's state-averaged effects and its 20,000-draw bootstrap with seed 107031. Then reconstruct PyBullet's bank-stratified primary analysis by averaging tasks and the three fixed checkpoint effects within state before 10,000 resamples. For a diagnostic control, deliberately use independent state-index vectors in checkpoint–bank cells while keeping paired cost differences and weights fixed. Point estimands should agree, whereas bootstrap distributions can differ. The correctly paired analysis should recover the reported V107 gate and negative PyBullet adjusted upper bounds within numerical precision; a change in point estimates signals an aggregation or alignment error rather than an uncertainty-only repair. e-estimandse-inferencee-v107e-primary-comparisonse-archive
8.3 Reading coverage
Visual audit: Visually inspected the title/author/version page and every page of the 20-page PDF, including all six figures, Tables 1–6, the proofs, Appendix C protocols, Appendix D comparisons, Appendix E fixed-threshold sensitivity, Appendix F guardrails and Appendix G archive disclosures. Six faithful final crops were separately inspected; table crop boundaries were refined to exclude caption fragments. Figure 1's routing branches and offline-only arrow were checked against the action definitions and Eqs. (1)–(2); its net-benefit label is distinguished from the gross fitting target in §3.1. All cited method, numerical and reproducibility pages are included. No separate supplement or release archive was supplied for this visual pass.
PDF pages inspected for this edition: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20. Appendix coverage: reviewed.
Original figures and tables remain the work of the source’s authors. Extractions preserve their scientific content; any HTML wrapper layout is disclosed with each figure. The surrounding reading notes are our own.
Text reading scope & known omissions
- Title, author block, version stamp and Abstract (p. 1)
- 1 Introduction (pp. 1–2)
- 2 Decision ledger, §§2.1–2.4 and Propositions 1–5 (pp. 2–5)
- 3 Paired physical routing, §§3.1–3.3 (pp. 5–7)
- 4 Frozen audit evaluations (pp. 7–8)
- 5 Results, §§5.1–5.4 (pp. 8–13)
- 6 Related work (p. 13)
- 7 Limitations and conclusion (pp. 13–14)
- A Proofs, §§A.1–A.9 (pp. 15–16)
- B Why latent distance is not a common currency (p. 16)
- C Protocol details, §§C.1–C.3 (pp. 16–17)
- D Complete primary comparisons, §§D.1–D.2 (p. 18)
- E V107 fixed-threshold price sensitivity (p. 18)
- F Claim guardrails and G Reproducibility and archival note (p. 19)
- References (pp. 19–20)
Outside the original text pass
- Text extraction does not reconstruct figure images; inspect the retained PDF for figures and equation/table layout.
- Separate supplemental material availability has not been fully verified.
- The inspected artifact is arXiv:2608.14650v1 [cs.LG], dated 30 July 2026. Its title matches the catalog; the observed author Malo de Pastor corresponds to the catalog's inverted string 'Pastor, Malo de'. No other revision or edition was supplied or compared.
- Separate supplemental material availability has not been fully verified; no supplement was supplied.
- Text extraction did not reconstruct images; this limitation was addressed by inspecting every PDF page, all numbered figures and tables, and every retained crop.
- The release archive, code, model checkpoints and raw data were not inspected. No paper scripts were executed and no experiment was reproduced.
The visual audit above records the subsequent illustrated pass.
8.4 Traceable evidence
e-identityPDF p. 1, title/author/affiliation block and left-margin arXiv stamp
The title matches the supplied identity. The author is Malo de Pastor, affiliated with Télécom SudParis, Institut Polytechnique de Paris, France. The artifact displays arXiv:2608.14650v1 [cs.LG], 30 July 2026.
Go to primary source ↓e-problemPDF pp. 1–2, Abstract, §1 question and contributions; p. 5, Table 1
The stated contribution is paired physical evaluation and audit of separately frozen capacities; standard adaptive-routing theory organizes the cost ledger.
Go to primary source ↓e-diagramPDF p. 2, Figure 1, boxes, arrows and caption; p. 6, §3.1 fitted-target paragraph
The diagram separates deployment from offline exact-reset labels and labels B as net benefit. The executed regressors instead fit gross regret difference and use validation-selected strict thresholds.
Go to primary source ↓e-pairingPDF pp. 2–3, §2.1 action/regret definitions and reset/replay paragraph; p. 5, §3.1
Shared candidate actions and exact resets reveal every realized candidate outcome. Regret subtracts the best realized candidate loss. Duplicate replay validation samples three states per PushT shard and one candidate per checked state.
Go to primary source ↓e-ledgerPDF p. 3, §2.2 Eqs. (1)–(4), Propositions 1–2; p. 15, §§A.1–A.2
Conditional net escalation benefit determines the Bayes rule; standalone Full avoids the sequential Medium/interface pass, producing the signed latency-gap correction.
Go to primary source ↓e-theoryPDF p. 4, Theorem 1, Proposition 3 and Corollary 1; p. 5, Table 1; p. 15, §§A.3–A.5
Nested information has a nonnegative population positive-part value gap. Finite learners can incur weighted sign regret, separate from acquisition overhead.
Go to primary source ↓e-routerPDF p. 6, §3.1 gross-target fitting, threshold selection and inference paragraphs
MLPs with 128/64 ReLU hidden units regress gross regret improvement. Validation selects strict thresholds from 101 score quantiles plus constant endpoints. Incremental latency is constant within evaluated checkpoint pairs.
Go to primary source ↓e-controlPDF p. 6, §3.2, interface dimensions and prospective seal
V107 uses dimension-matched 65-D prediction and current-DINO/action inputs; preprocessing is fit only on V106 training states. The control routes before Medium and pays no DINO encoder latency. Fresh outcomes follow freezing the comparison and artifacts.
Go to primary source ↓e-restrictionsPDF p. 16, Appendix C.1
Prediction routers exclude explicit simulator state and privileged realized outcomes; PushT uses task and calibrated prediction, while PyBullet additionally uses a prespecified regime. The current-DINO/action control is a declared input-routing exception.
Go to primary source ↓e-pushtPDF p. 7, §4 PushT and V107 paragraphs; p. 8, Table 3, PushT rows
PushT supplies five candidates at horizon 30 and 39 target/weighting tasks. Predictor architectures, AdamW row counts, ridge calibration, 5,200/1,200 router state split, and separate 1,600-state V106 and V107 audits are specified.
Go to primary source ↓e-pybulletPDF pp. 7–8, §4 Controlled PyBullet; p. 8, Table 3, PyBullet rows
Six 2,000-state banks and 81 tasks evaluate three checkpoint pairs. Residual DINO-feature Transformers use 192/4/2 and 384/6/3 width/heads/layers; training is 100 AdamW epochs on 5,000 states, with 1,000 calibration states. Parameter/MAC counts are not supplied.
Go to primary source ↓e-timingPDF p. 8, §4 Latency protocol
The decision unit is one state, one task and five candidates. H100 forwards and single-threaded CPU components are measured separately; PushT uses 200 warmups, seven 1,000-repeat rounds and 512 queries. CPU SKU is missing.
Go to primary source ↓e-inferencePDF p. 7, §3.3 bootstrap repair and comparison-family paragraphs
The original PyBullet bootstrap failed to preserve shared states across checkpoint positions. Repaired primary inference resamples physical states within fixed banks after averaging fixed checkpoints, without changing point estimates or policies.
Go to primary source ↓e-estimandsPDF p. 17, Appendix C.2–C.3, gates, estimands and bootstrap definitions
V107 averages tasks/checkpoints within 1,600 states, with 20,000 bootstrap draws and seed 107031. V106 uses a 0.99 adjusted upper quantile; PyBullet uses 10,000 draws within six banks and a 0.9875 upper quantile. Crossed sensitivities are not checkpoint-population intervals.
Go to primary source ↓e-reversalsPDF p. 8, Figure 2, legend/caption and §5.1
After task and checkpoint averaging, Medium has lower decision cost on 45.1% of PushT and 51.9% of PyBullet states; Full wins on the complements. Realized heterogeneity alone does not prove exploitable interface value.
Go to primary source ↓e-v107PDF p. 9, Figure 3 and §5.2 first paragraph; p. 18, Appendix D.1 all rows
The aggregate prediction-minus-control effect is −0.002549 with 95% interval [−0.002867, −0.002238] and upper bound −0.002286. Fixed checkpoint effects for 92002, 2026 and 314159 are −0.002680, −0.002428 and −0.002540.
Go to primary source ↓e-v107-tablePDF p. 9, Table 4 all rows and following interpretation paragraph
Prediction/control regret is 0.02849/0.03110, latency 0.664/0.630 ms, priced cost 0.02981/0.03236 and escalation 47.2%/66.2%. The text reports a 0.002617 regret gain, 0.000068 priced latency disadvantage and 7.9% cost reduction.
Go to primary source ↓e-v106PDF p. 9, §5.2 Earlier V106 comparison; p. 10, opening benchmark paragraph
Task-only escalation averages 99.1453%; prediction-minus-task-only cost is −0.002620 with 95% interval [−0.003037, −0.002213]. Margin/random differences and a descriptive 1/5001 matched-count tail fraction are reported.
Go to primary source ↓e-primary-comparisonsPDF p. 10, Figure 4 and §§5.2–5.3; p. 18, Appendix D.2
V106 fixed-Medium/Full effects are −0.004308/−0.002577. PyBullet effects are −0.003677/−0.003746 versus fixed capacities and −0.002800/−0.003668 versus margin/random, with repaired intervals. Adjusted primary upper bounds remain negative.
Go to primary source ↓e-operatingPDF p. 11, Table 5 all rows, Figure 5 caption and §5.4
PushT router regret/latency/cost is 0.02907/0.651 ms/0.03037; PyBullet is 0.02309/1.320 ms/0.02573. Both fixed policies are faster. The clairvoyant lower bound uses realized outcomes and zero router overhead.
Go to primary source ↓e-pricePDF p. 12, Figure 6 and Table 6 all rows; p. 13, §5.4 continuation
V106 thresholds are validation-selected for each declared compute price. Differences favor routing through the tested price 0.01, but favor every listed reference at 0.025 and 0.05; Full difference at 0.01 is −0.000385.
Go to primary source ↓e-v107-pricePDF p. 18, Appendix E text and table
V107 holds its lambda=0.002 routers/thresholds fixed. Prediction-minus-control stays negative at all six declared prices from 0 to 0.05, without claiming dominance over fixed policies.
Go to primary source ↓e-replayPDF p. 13, §5.4 threshold-replay paragraph
Exact serialized thresholds and NumPy float32 with strict score>threshold are scientific replay semantics. A parser/promotion change alters 6,400 high-price decisions among a grid containing 17,600 score–threshold equalities.
Go to primary source ↓e-recodingPDF p. 16, Appendix B, Proposition 7 and explanation
Invertible measurable finite-support recoding preserves unrestricted decision information while changing coordinate geometry; it need not preserve neural learnability or compute cost.
Go to primary source ↓e-limitsPDF pp. 13–14, §7; p. 19, Appendix F
The source limits inference to related DINO-style models, fixed checkpoint pairs and controlled simulator tasks. It disclaims closed-loop value, universal prediction superiority, compute saving and synchronized end-to-end timing evidence.
Go to primary source ↓e-archivePDF p. 19, Appendix G, release layers, omissions and tie paragraph
The described compact release supports statistics and router-lineage checks, but omits raw observations, complete candidate outcomes, feature banks, upstream models and original HPC environment. Seals lack external chronology certification; action ties use first-minimum argmin.
Go to primary source ↓8.5 Primary sources
Paired Exact-Reset Evaluation of a Prediction-Derived Medium-to-Full World-Model Cascade ↗
PDF · 9,626 extracted words
Source fingerprint
b5811b941c7dc0584b96d9462a826939c873988ff8acce3e795dc389ec30cd8b