PAPER REPORTENAll readings ↗

Dream-Tac: A Unified Tactile World Action Model for Contact-Rich Robot Manipulation

English reading report: Method, equations, original figures, experiments and reproducibility.

Authors: Yunfan Lou; Yifan Ye; Yankai Fu; Jun Cen; Xiaowei Chi; Yaoxu Lyu; Peidong Jia; Sirui Han; Zhihe Lu; Shanghang Zhang

Affiliations: Peking University, Beijing, China; The Hong Kong University of Science and Technology, Hong Kong, China; Nanjing University, Nanjing, China; State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University, Beijing, China

Source: 2606.08737 ↗ · Catalog record

Reading: 179 / 558 · 6 original figures & tables · ~19 min ·

1. Paper overview

In one sentence: Dream-Tac jointly denoises actions and visuo-tactile futures, using a deterministic contact-change bias to improve real robot success at the cost of an enlarged diffusion model. e-motivatione-jointe-casae-maine-ablatione-efficiencye-gate-diagnostice-reconstruction

At a glanceWhat to know
Research problem
Source description

Vision can leave contact onset, slip and local deformation ambiguous. Touch supplies these cues, but sparse events coexist with long, nearly static intervals. Dream-Tac asks how a world action model can use tactile observations selectively while jointly anticipating future interaction and generating executable action chunks. e-motivatione-joint

Core mechanism
Source description

A shared latent sequence combines visual, tactile, state and action tokens within one diffusion transformer, reusing a pretrained video backbone and its text encoder/VAE. e-jointe-architecture

A key reported resultSix-task real-world manipulation: 83.3

Mean task success rate (%). Franka Panda, two D435i cameras and two Xense Photon tactile sensors; 100 demonstrations per task; 20 evaluation trials per method/task, best checkpoint under a shared selection rule.

Cosmos Policy 51.7; ForceVLA 50.8; pi0.5 45.0; pi0 30.8. Displayed means differ by 31.6 percentage points; averaging the per-task gaps gives 31.7 points after rounding. The abstract calls this action accuracy, although the experiment measures task success. e-datae-protocole-maine-headline

Reading caution
Source description

The authors acknowledge limited task/environment coverage, a simple change-based gate that may miss subtle or long-horizon interaction, and continued diffusion cost relative to lightweight reactive policies. e-limitations

Core contributions

  • Source description

    A shared latent sequence combines visual, tactile, state and action tokens within one diffusion transformer, reusing a pretrained video backbone and its text encoder/VAE. e-jointe-architecture

  • Source description

    Contact-aware self-attention adds a deterministic, directed tactile bias; a rank-one attention reformulation and fixed diffusion-step cache reduce the expanded model’s computational cost. e-casae-gatee-flashe-cache

Figure 2. Actions and sensory futures share one denoising backbone. Original paper, p. 4 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Read the left side from bottom to top. Visual and tactile images pass through a VAE, while robot state and action enter as latent-frame tokens. Solid tokens identify clean context; hatched future/action tokens are denoising targets. The shared Dream-Tac block produces action latents alongside future sensory latents, with optional image decoding. On the right, text embeddings enter DiT cross-attention, while CASA modifies self-attention. The lower inset illustrates factorized bias channels. Use Eq. (6) to determine its direction: non-tactile queries receive extra access to tactile keys. The generic Q/K labels in the drawing should not be read as a symmetric tactile bias. e-jointe-architecturee-casae-gatee-objectivee-training

What it supports. The architecture supports the catalog’s One Model × Joint prediction assessment: future vision, future touch and actions interact inside the same DiT. Touch is both an observed input and a generated target. Optional image decoding does not mean the future latent tokens disappear from the joint generative process.

Where the evidence stops. The gate inset labels VT_t and VT_(t+1), whereas Eq. (7) uses observed t−1 and t. CASA’s tactile/other colors also differ from the left panel. Follow the equations’ observed-only, directed formulation; the schematic does not resolve gate broadcasting over the joint sequence.

2. Motivation

2.1 The problem and the proposed response

Source description

Vision can leave contact onset, slip and local deformation ambiguous. Touch supplies these cues, but sparse events coexist with long, nearly static intervals. Dream-Tac asks how a world action model can use tactile observations selectively while jointly anticipating future interaction and generating executable action chunks. e-motivatione-joint

2.2 What this reading follows

A robot can see a tool near an object without knowing whether useful contact has begun. Dream-Tac brings fingertip images into the same predictive model that generates robot actions, so anticipated touch can interact with anticipated visual motion. Its distinctive intervention is a small, directed attention bias controlled by changes in observed tactile images. The experiments separate adding tactile modeling from adding this bias, and an engineering study examines the cost of both. Read the six visuals as complementary evidence: an architectural specification, executed task outcomes, a staged ablation, timing measurements, a gate diagnostic and qualitative tactile predictions. Their limitations differ, and no single panel validates the whole causal story. e-motivatione-jointe-casae-maine-ablatione-efficiencye-gate-diagnostice-reconstruction

3. Research context

We place the paper in the collection through its world–action interface. The catalog labels and the reading’s assessment are shown separately.

Catalog dimensionRecorded classification
Major categoryWAMs
ArchitectureOne Model
Prediction paradigmJoint prediction
QuadrantQ1 · One Model × Joint prediction

3.1 Evidence-based assessment

Supports the recorded classification

Reader analysis

One Model × Joint prediction is supported architecturally: the same DiT jointly denoises action, visual-future and tactile-future tokens with bidirectional attention. T5/VAE encoders do not constitute a separate world-model-to-policy cascade. Tactile and efficiency tags are supported; no audio experiment is presented despite the broader catalog subcategory wording. e-jointe-architecturee-objectivee-efficiency

This is the collection’s architectural analysis, not a new related-work survey. Benchmark comparisons and their protocols appear in Section 6.

4. Problem formulation

4.1 Inputs and outputs

InputsOutputs
  • Current third-person and wrist RGB observations
  • Left/right fingertip tactile RGB and preceding observed tactile frames for the gate
  • Robot proprioceptive state
  • Language instruction
  • A jointly denoised action chunk with H = 20
  • Future visual and tactile latents, optionally decoded into images

4.2 Equations and their role

p(a1:H,v1:T,x1:To,x,l)p(a_{1:H},v_{1:T},x_{1:T}\mid o,x,l)
Equation (5): jointly model action chunk a over H steps and visual/tactile futures v/x over T steps, conditioned on current visual observation o, tactile observation x and instruction l. e-joint
logitij=qikjd+αgt(1Mi)Mj\operatorname{logit}_{ij}=\frac{q_i^\top k_j}{\sqrt d}+\alpha g_t(1-M_i)M_j
Equation (6): q/k are query/key vectors, d is head dimension, M marks tactile tokens, and g is the tactile-change gate. The addition applies only to non-tactile queries and tactile keys; alpha = 2.0 in Appendix A.5. e-casae-training
ρt=max(δtL,δtR),zt=kρtms+ϵ,gt=gmin+(gmaxgmin)σ(zt)\rho_t=\max(\delta_t^L,\delta_t^R),\quad z_t=k\frac{\rho_t-m}{s+\epsilon},\quad g_t=g_{\min}+(g_{\max}-g_{\min})\sigma(z_t)
Equations (7)–(9): delta is normalized mean absolute image change per fingertip, rho their maximum, and sigma the sigmoid. Fixed (m,s,k,epsilon) = (0.002,0.001,4,10^-6); z is clipped to [-30,30]. Gate bounds [0.15,1.0] reduce, but never eliminate, the bias. e-gate
y~=y+σϵ,Ldenoise=Ey,ϵ,σ[fθ(y~,σ,o,x,l)ϵ22]\widetilde y=y+\sigma\epsilon,\qquad\mathcal L_{\mathrm{denoise}}=\mathbb E_{y,\epsilon,\sigma}\left[\left\|f_\theta(\widetilde y,\sigma,o,x,l)-\epsilon\right\|_2^2\right]
Equations (10)–(11): y contains future visual, tactile and action latents; epsilon is Gaussian noise, sigma the sampled noise level, and f_theta the backbone. This is the main-text presentation, subject to the appendix parameterization caveat above. e-objectivee-training

5. Method in detail

5.1 Follow observed touch and predicted touch through different roles

Source description

Start with the joint distribution in Eq. (5). Current visual and tactile observations condition a joint sample of robot actions and future sensory observations. This distinction matters when reading Figure 2: current sensor tokens form clean context, whereas action and future tokens are jointly corrupted during training and jointly denoised. The same VAE maps camera and tactile RGB into the backbone’s latent space, and padded latent-frame tokens accommodate state and actions. Language enters through cross-attention. Bidirectional self-attention allows the action tokens to interact with sensory futures inside one model. Thus future touch is an explicit prediction target, while observed touch is conditioning information. Optional decoding exposes the predicted sensory latents as images; it is not a separate inverse-dynamics stage that converts a completed video into actions. e-jointe-architecturee-objectivee-training

5.2 Separate the tactile-change heuristic from learned attention

Reader analysis

The gate does not learn a contact detector. Equations (7)–(9) transform consecutive observed tactile images into a scalar: average absolute pixel change per fingertip, take the larger change, normalize with fixed constants, and apply a bounded sigmoid. Equation (6) then uses that scalar only to favor tactile keys for non-tactile queries. Learned query–key content still determines which tokens are useful. Reader interpretation: this separates a coarse event heuristic from fine attention allocation, but it also makes sensor disturbances a plausible failure mode. Figure 7 shows both sustained activation and sharp spikes without contact ground truth. Moreover, the positive gate floor keeps some bias active during quiet periods. A comparison against constant bias is therefore necessary to attribute the ablation gain specifically to temporal selectivity. e-gatee-casae-gate-diagnostice-ablation

5.3 Ask which computation can be removed without changing the claim

Reader analysis

The two accelerations have different scientific status. Appendix B.1 algebraically rewrites the rank-one bias into extra query/key channels; preserving the intended scaling should preserve the attention logits. Appendix B.2 instead reuses intermediate computation between diffusion steps, which is an approximation. Its validation diagnostics show that action changes are poorly tracked by timestep embeddings, motivating fixed full evaluations at the first and third steps. Figure 5 tests this choice on Peel Cucumber and reports equal success with lower latency. Reader interpretation: the fused-attention change invites a numerical equivalence check, whereas caching requires both prediction-error and closed-loop control checks. Similarity between adjacent latents does not guarantee that a small difference is harmless near insertion or contact transitions. Neither efficiency result establishes universal real-time performance across hardware and tasks. e-flashe-cachee-efficiencye-timing-contexte-main

5.4 Training and inference

During training

Source description

Appendix A.5 fine-tunes Cosmos-Predict2-2B Video2World with Fused Adam, learning rate 10^-4, betas (0.9, 0.99), weight decay 0.1 and bfloat16. RGB is 224 × 224; states/actions are standardized. Chunk length is 20; per-GPU batches are 25 for vision-only and 16 for visuo-tactile layouts. A.4 describes eight H100 GPUs for Cosmos Policy under the same configuration as Dream-Tac. e-training

Reader analysis

Observed context stays clean while future/action tokens are jointly noised. Section 3.4 presents Gaussian corruption and an epsilon-prediction MSE with modality losses. Appendix A.5 specifies the inherited rectified-flow/hybrid-EDM objective and hybrid noise sampling. Their exact correspondence, modality weights and frozen-module policy are not fully specified. e-objectivee-training

During inference

Reader analysis

Condition on current observations, state and instruction, jointly denoise action/future tokens, and read out the action chunk; image decoding is optional. The gate uses observed tactile history. The paper does not fully specify how many chunk actions execute before observation refresh or the low-level control interface. e-jointe-architecturee-gatee-training

Source description

For ten-step cached inference, full forward computation occurs at the first and third steps; remaining steps reuse cached computation. This retains a ten-step schedule with two full evaluations. e-cachee-efficiency

5.5 Implementation flow

  1. Build a shared latent sequence

    The pretrained VAE encodes visual and tactile images. Robot states and actions enter as padded latent-frame tokens. T5 language embeddings reach DiT blocks through cross-attention. Bidirectional self-attention lets action tokens interact with jointly predicted visual and tactile futures; a separate inverse-dynamics decoder is not described. e-jointe-architecturee-training

  2. Measure tactile change and bias attention

    Compute each fingertip’s mean absolute RGB change from the preceding observed frame, normalize by 255, then take the larger value. Fixed normalization and a sigmoid produce a bounded gate without a learned gating network. CASA adds bias only from non-tactile queries to tactile keys; content attention still chooses individual tokens. e-casae-gatee-gate-diagnostic

  3. Preserve fused attention

    The directed bias is rank one. Appendix B.1 appends scalar channels to queries and keys so their dot product equals the original scaled content score plus bias. Zero padding accommodates kernel alignment without materializing a dense bias matrix. e-flash

6. Experiments & results

Dream-Tac fine-tunes a shared video diffusion transformer to generate robot actions together with future camera and tactile latents. A deterministic tactile-change gate directs attention toward touch, while fused attention and diffusion caching reduce computational cost. Real robot experiments support the multimodal design, but do not isolate the benefit of predicting future touch from simply observing it.

6.1 Read the original evidence

Figure 3. The mean improvement includes large gains and one difficult insertion task. Original paper, p. 7 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Read the six task panels before the final Average panel. Each group uses the same five-method order, and the hatched bar is Dream-Tac. The y-axis is task success percentage, not action-coordinate error or reconstruction quality. The caption specifies twenty real-world trials per method and task, using the best checkpoint under a shared selection rule. Compare the matched baguette bars, the low USB insertion bars and ForceVLA’s stronger cucumber result before interpreting the overall mean. The paper describes Mahjong as visually occluded, so its success should be read in that particular sensing condition rather than as general visual game understanding. e-maine-protocole-headlinee-training

What it supports. Dream-Tac averages 83.3% against Cosmos Policy’s 51.7%, a 31.6 percentage-point difference between displayed means. Its USB result is 35% versus 15%, while Cut Banana rises from 40% to 90%. Pick Baguette ties at 100%; ForceVLA exceeds Dream-Tac on Peel Cucumber, 90% versus 85%.

Where the evidence stops. Twenty trials per task and best-checkpoint reporting give limited precision; no uncertainty intervals are shown. The abstract’s 31.7 figure follows rounding the per-task gap, but the metric is task success and the gain is in percentage points. Insertion remains unreliable.

Figure 5. Fused bias computation and cached denoising address different costs. Original paper, p. 8 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Treat the two panels as separate experiments on Peel Cucumber. On the left, compare the two methods within each tactile/bias configuration; lower seconds are better. The largest difference occurs when both touch and bias are enabled. On the right, blue bars report success, purple bars carry explicit millisecond labels, and the green line uses the right-hand success-to-latency axis. Read latency from its numerical annotations, not the success-rate ticks. The last condition is labeled ten steps with cache. Appendix B.2 explains that full computation occurs only at the first and third steps while cached computation is reused elsewhere. e-efficiencye-timing-contexte-flashe-cache

What it supports. In the full tactile-plus-bias configuration, training time falls from 80.82 to 27.48 seconds, approximately 2.9× faster. Cached ten-step inference reports 619 ms instead of 1109 ms with the same 85% success on this task. Simply using one step reaches 481 ms but lowers success to 60%.

Where the evidence stops. The figure does not specify the training-time aggregation or clearly attach timing hardware. Section 3.5 separately reports H200 iteration times and A800 frequency; those contexts are not reconciled here. Equal observed success on one task does not prove cache equivalence.

Figure 13. Predicted touch reflects a qualitative transition from pre-contact to cutting. Original paper, p. 16 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Read across each row before comparing rows. The left image is ground truth and the right image is prediction; the upper pair shows pre-contact and the lower pair cutting on Cut Banana. The red rectangles belong to the original figure and highlight the lower contact-sensitive regions. Compare broad deformation and intensity patterns there, rather than treating every pin as a calibrated force measurement. Then inspect whether the phase change from the upper to lower row appears in both columns. Appendix C.2 interprets this as phase-dependent tactile prediction and acknowledges smoothing of finer pin-level detail. e-reconstructione-jointe-ablation

What it supports. The example is consistent with the intended predictive target: Dream-Tac can produce a relatively undeformed pre-contact pattern and a different cutting-phase pattern. This complements the robot-success results by making future touch visible. It supports qualitative correspondence in these examples, without assigning an error magnitude or establishing general prediction reliability.

Where the evidence stops. The source supplies selected examples without a reconstruction-error metric, force calibration or matched prediction baseline. Fine detail is visibly smoothed. These images alone cannot establish physically accurate forces or show that the future-tactile objective caused the policy’s success gains.

6.2 Results and evaluation conditions

Task & protocolReported resultComparison & interpretation
Six-task real-world manipulation

Franka Panda, two D435i cameras and two Xense Photon tactile sensors; 100 demonstrations per task; 20 evaluation trials per method/task, best checkpoint under a shared selection rule.

83.3

Mean task success rate (%)

Cosmos Policy 51.7; ForceVLA 50.8; pi0.5 45.0; pi0 30.8.

Displayed means differ by 31.6 percentage points; averaging the per-task gaps gives 31.7 points after rounding. The abstract calls this action accuracy, although the experiment measures task success. e-datae-protocole-maine-headline

Per-task real-world manipulation

Same six-task evaluation; Figure 3 task order.

Pick Baguette 100; Insert USB 35; Clean Whiteboard 90; Peel Cucumber 85; Play Mahjong 100; Cut Banana 90.

Success rate (%)

Cosmos Policy: 100, 15, 55, 65, 35, 40 respectively. ForceVLA reaches 90 on Peel Cucumber.

Dream-Tac ties the best baguette result and leads four other tasks, but USB insertion remains unreliable. Mahjong is described as visually occluded. e-maine-protocol

Tactile fusion and CASA ablation

Table 1, mean over six real-world tasks.

Visual WAM 51.7; visuo-tactile WAM 74.2; visuo-tactile WAM + bias 83.3.

Average success rate (%)

Adding tactile modeling gains 22.5 points; adding bias gains another 9.1 points, calculated from displayed means.

The staged comparison supports tactile modeling and bias, but provides no constant-bias control or input-only tactile variant. e-ablation

Environment-variation generalization

Figure 4: height shifts for Peel Cucumber, unseen placements for Pick Baguette, unseen tiles for Play Mahjong, changed background for Cut Banana.

Dream-Tac: +5/-5 cm height 90/75; OOD placement 80; unseen tiles 85; altered background 70.

Success rate (%)

Cosmos Policy: 30/0; 80; 15; 25, respectively.

Advantages depend on the perturbation; spatial generalization ties. These are selected task/variation pairs, not broad cross-task transfer. e-ood

Peel Cucumber computational efficiency

Figure 5: full tactile-plus-bias training configuration; ten-step inference with or without caching.

Training 27.48 s; cached inference 619 ms at 85% success.

Training time (s), inference latency (ms), task success (%)

Training reference 80.82 s; uncached inference 1109 ms at 85%.

Approximately 2.9× training and 1.8× inference speedups for these comparisons. Figure 5 does not clearly identify training-time aggregation or timing hardware; other timing claims should not be merged with it. e-efficiencye-timing-context

6.3 Ablations and diagnostic examples

Read component removals and qualitative examples within their stated evaluation conditions.

Table 1. Adding tactile modeling supplies the larger gain; CASA adds another increment. Original paper, p. 6 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Read downward as a staged intervention. The first row has neither tactile input nor attention bias. The second introduces tactile modeling but retains ordinary self-attention. The third retains touch and activates the proposed bias. The rightmost column averages success over the same six real-world tasks. This structure distinguishes the combined tactile-modeling addition from the subsequent CASA addition. It does not separate the roles of observing present touch and predicting future touch: both belong to the visuo-tactile design. Likewise, the final switch adds the proposed gated bias without comparing it against an equally strong bias that stays constant in time. e-ablatione-jointe-traininge-protocol

What it supports. The reported means rise from 51.7% to 74.2% to 83.3%. Subtracting displayed values gives a 22.5-point gain for tactile modeling and a further 9.1-point gain for CASA. These results support the complete additions under the reported protocol, while leaving their internal mechanisms open to more controlled testing.

Where the evidence stops. There are no per-task ablation scores, repeated-seed intervals, input-only tactile controls or constant-bias controls here. The training appendix also uses different batch sizes for vision-only and visuo-tactile layouts, so identical training resources should not be assumed.

Figure 7. The gate responds to tactile image change, including brief spikes. Original paper, p. 8 ↗

Excerpt from the authors’ paper; cropped without altering the figure or table.

How to read it. Each row is one randomly sampled Peel Cucumber training demonstration. Follow blue tactile change on the left axis and red gate on the right; their units differ. The dotted horizontal line marks that episode’s 75th percentile of tactile change, not a learned decision boundary or the gate’s sigmoid midpoint. The gate is obtained deterministically from consecutive observed fingertip images using the maximum left/right change and a fixed sigmoid mapping. Compare long elevated intervals with the thin spikes in quieter intervals. The five-episode analysis includes 874 noninitial timesteps; Appendix A.6 supplies the gate’s lower and upper bounds. e-gate-diagnostice-gatee-casa

What it supports. The traces show that the gate is not constant: changes in tactile appearance can move it across much of its allowed range. The authors associate sustained elevation with contact-rich phases and lower intervals with approach or release. The diagnostic illustrates the input-to-gate mapping used by CASA, rather than measuring downstream action improvement.

Where the evidence stops. Several brief spikes reach high gate values. The paper attributes approach oscillations partly to sensor noise or disturbances, but supplies no independent contact labels here. A deterministic response to image change is not a measured contact-detection precision or recall result.

7. Analysis & limitations

7.1 What the evidence leaves open

Source description

The authors acknowledge limited task/environment coverage, a simple change-based gate that may miss subtle or long-horizon interaction, and continued diffusion cost relative to lightweight reactive policies. e-limitations

Reader analysis

No uncertainty intervals or repeated-seed statistics accompany success means. Best-checkpoint reporting, an unspecified validation split, and differing baseline compute/action horizons limit causal comparison. Qualitative reconstructions and t-SNE do not establish calibrated contact dynamics or prove that future tactile prediction improves control. e-protocole-traininge-ablatione-reconstructione-latent

Reader analysis

Source discrepancies remain: Figure 2 labels gate images t and t+1, but Eq. (7) uses observed t-1 and t. Section 4.1 requires peeling/cutting thresholds exceeding 5 cm, while Appendix A.2 uses qualitative success descriptions. e-architecturee-gatee-task-definitions

7.2 Questions for discussion

  1. Does predicting future tactile latents improve control when current tactile input and training compute are fixed?
  2. Does the changing gate outperform constant bias of matched average magnitude under sensor noise and sustained static contact?
  3. How robust is the first/third-step cache schedule when contact dynamics or task horizons change?

8. Reproducibility audit

8.1 Requirements and known gaps

Reader analysis

Reproduction requires synchronized camera/tactile/proprioceptive demonstrations collected at 30 Hz, the specified robot/sensors and pretrained checkpoint. Table 2 supplies execution limits. Missing details include split membership, seeds, total Dream-Tac training steps, modality weights, future horizon T, frozen modules and precise cache tensors. e-datae-traininge-objectivee-cache

Reader analysis

Reader-proposed priorities are matched input-only versus future-tactile prediction and constant-bias versus changing-gate controls. Verify dense/augmented attention equivalence and benchmark cache errors on held-out contact transitions before interpreting speedups as preserved control quality. e-ablatione-flashe-cachee-gate

8.2 Proposed reproduction checks

The following checks are proposals motivated by the paper. They have not been run as part of this reading.

Check 1: Separate present touch, predicted touch and a changing gate

Reader-proposed experiment, not performed: use the same demonstration split, initialization, action horizon and optimizer budget for four models: current tactile input without a future-tactile loss; joint tactile prediction without bias; joint prediction with constant bias; and joint prediction with the observed-change gate. Keep future-token layout and token count matched where possible, and disclose the exact treatment of unsupervised future tokens. Match constant-bias magnitude to the training-set mean gate, never to test episodes. Evaluate fixed held-out initial states on USB insertion and cucumber peeling, with repeated training seeds. If future supervision does not improve over input-only sensing, or the changing gate does not beat constant bias, the corresponding causal interpretation of Table 1 is weakened. e-ablatione-jointe-objectivee-gatee-traininge-protocol

Check 2: Verify exact attention before testing approximate cache reuse

Reader-proposed two-stage check, not performed: first compare dense Eq. (6) attention against Appendix B.1’s augmented queries/keys using identical tensors, token masks and gate extrema. Compare logits, outputs and gradients in float32 and bfloat16; ensure the fused operator does not apply an unintended second scaling. Then compare full ten-step denoising with first/third-step caching using identical starting noise and held-out episodes. Record action-latent error by contact phase, synchronized end-to-end latency on one specified GPU, and robot success with the same checkpoint and trial definitions. A cache that saves time but increases failures near contact would falsify transferring the cucumber result to other tasks, even if average latent similarity remains high. e-casae-gatee-flashe-cachee-efficiencye-protocole-task-definitions

8.3 Reading coverage

Visual audit: All sixteen supplied PDF pages were rendered and visually inspected, including the title/author block, Figures 1–13, Tables 1–2, equations, main experiments, task definitions and Appendices A–D. All six final crops were separately viewed; the efficiency crop was widened to retain its right-axis title and the tactile crop adjusted to retain labels and borders. Numerical, method, training, hardware, evaluation and proposed-check evidence pages are included even where not cropped. Figure 2’s temporal labels and CASA colors are disclosed against the equations; independent contact labels, quantitative reconstruction errors and full implementation details are not supplied. No linked code, separate supplement or later version was inspected.

PDF pages inspected for this edition: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16. Appendix coverage: reviewed.

Original figures and tables remain the work of the source’s authors. Extractions preserve their scientific content; any HTML wrapper layout is disclosed with each figure. The surrounding reading notes are our own.

Text reading scope & known omissions
  • Abstract
  • 1 Introduction
  • 2 Related Work
  • 2.1 Tactile in Robot Manipulation
  • 2.2 World Action Model
  • 3 Method
  • 3.1 Problem Formulation
  • 3.2 Dream-Tac Architecture
  • 3.3 Contact-Aware Self Attention
  • 3.4 Training Objective
  • 3.5 Efficiency Design
  • 4 Experiment
  • 4.1 Experimental Setups
  • 4.2 Evaluations
  • 4.3 Training and Inference Efficiency
  • 4.4 Analysis
  • 5 Conclusion
  • Acknowledgments
  • References
  • A Additional Experimental Details
  • A.1 Real-World Experimental Setup
  • A.2 Task Details
  • A.3 Experimental Settings
  • A.4 Baseline Settings
  • A.5 Training Hyperparameters
  • A.6 Contact Gate Statistics
  • B Implementation Details of the Acceleration
  • B.1 FlashBias Implementation
  • B.2 Diffusion-Step Time Cache
  • C Additional Reconstruction Results
  • C.1 Visual Reconstruction
  • C.2 Visuo-Tactile Reconstruction
  • D Limitations and Future Work

Outside the original text pass

  • Text extraction does not reconstruct figure images; inspect the retained PDF for figures and equation/table layout.
  • Separate supplemental material availability has not been fully verified.
  • Identity: the supplied title page identifies arXiv:2606.08737v1 [cs.RO], 7 June 2026. Its exact title and all ten authors agree with the catalog; no different edition or revision was supplied for comparison.
  • Acquisition caveat preserved: text extraction does not reconstruct figure images or equation/table layout. This reading additionally inspected all sixteen original PDF pages, Figures 1–13 and Tables 1–2.
  • Separate supplemental material availability has not been fully verified; none was supplied.
  • Code and linked resources were not inspected, and experiments were not reproduced.

The visual audit above records the subsequent illustrated pass.

8.4 Traceable evidence

e-identityPDF p. 1, title, author/affiliation block and arXiv marginInspect

Exact observed title matches the supplied primary title; arXiv:2606.08737v1, 7 June 2026. Authors: Yunfan Lou, Yifan Ye, Yankai Fu, Jun Cen, Xiaowei Chi, Yaoxu Lyu, Peidong Jia, Sirui Han, Zhihe Lu and Shanghang Zhang. Affiliations identify Peking University, HKUST and Nanjing University, including Zhang’s Peking University laboratory/school.

Go to primary source ↓
e-motivationPDF p. 2, Section 1, contact ambiguity and transient tactile signalsInspect

Vision can miss contact cues; uniform attention can dilute sparse tactile interaction events.

Go to primary source ↓
e-jointPDF p. 3, Sections 3.1–3.2, Eq. (5) and architecture paragraphsInspect

The joint distribution covers actions and visual/tactile futures. A shared DiT uses T5 cross-attention, VAE image/tactile tokens and padded state/action frames with bidirectional self-attention.

Go to primary source ↓
e-architecturePDF p. 4, Figure 2 and captionInspect

The diagram distinguishes clean/noisy tokens, joint denoising and optional image decoding. Its gate inset labels VT_t and VT_(t+1); CASA token colors/labels differ from modality colors on the left, so directional claims require Eq. (6).

Go to primary source ↓
e-casaPDF p. 4, Section 3.3, Eq. (6) and following paragraphInspect

The additive alpha*g_t*(1-M_i)*M_j term biases non-tactile queries toward tactile keys; tactile-to-tactile attention is unchanged.

Go to primary source ↓
e-gatePDF p. 5, Eqs. (7)–(9) and observed-input paragraph; p. 13, Appendix A.6Inspect

Gate inputs are consecutive observed left/right tactile RGB frames. Fixed parameters are (0.002,0.001,4,10^-6), clipping is [-30,30], rho_0=0, and Appendix A.6 gives gate bounds [0.15,1.0].

Go to primary source ↓
e-objectivePDF p. 5, Section 3.4, Eqs. (10)–(12)Inspect

The main text describes Gaussian corruption, epsilon-prediction MSE and an action/image/tactile loss decomposition with unspecified lambda_v and lambda_t values.

Go to primary source ↓
e-trainingPDF p. 13, Appendices A.4–A.5Inspect

A.4 gives baseline compute and training settings, including eight H100s for Cosmos Policy under the same configuration as Dream-Tac. A.5 names Cosmos-Predict2-2B Video2World, optimizer/precision, hybrid noise objective, clean context, RGB resolution, normalized actions/state, chunk length 20, batches 25/16 and alpha=2.0. Validation-based selection is stated without split membership.

Go to primary source ↓
e-dataPDF p. 11, Appendices A.1/A.3 and Figure 8; p. 12, Appendix A.3; p. 13, Table 2Inspect

Franka Panda uses fixed/wrist D435i cameras and two Xense Photon sensors. Each task has 100 SpaceMouse demonstrations recorded at 30 Hz. Evaluation randomizes positions within an unspecified predefined range, uses 20 trials and task-specific maximum steps; Table 2 lists those limits.

Go to primary source ↓
e-protocolPDF pp. 5–6, Section 4.1; p. 6, Section 4.2.1; p. 7, Figure 3 captionInspect

Hardware and common success criteria are described; 20 trials per method/task and shared best-checkpoint selection are stated. Mahjong is described as fully visually occluded. No uncertainty intervals accompany the results.

Go to primary source ↓
e-mainPDF p. 7, Figure 3, all task panels and Average panel; p. 6, Section 4.2.1Inspect

Dream-Tac task success is 100/35/90/85/100/90, mean 83.3. Cosmos Policy is 100/15/55/65/35/40, mean 51.7. Other means are 50.8,45.0,30.8; ForceVLA scores 90 on Peel Cucumber.

Go to primary source ↓
e-headlinePDF p. 2, abstract continuation and Section 1; p. 1, Figure 1(c); p. 7, Figure 3Inspect

The abstract says 31.7% action accuracy improvement; the introduction and overview say 31.6%. Figure 3’s displayed means are 83.3% and 51.7% success.

Go to primary source ↓
e-ablationPDF p. 6, Table 1, all three rows; Section 4.2.2Inspect

Visual WAM, visuo-tactile WAM and visuo-tactile WAM plus bias score 51.7,74.2,83.3. The table does not separate tactile input from future-tactile supervision or compare dynamic and constant bias.

Go to primary source ↓
e-oodPDF pp. 6–7, Section 4.2.3; p. 8, Figure 4(a–d)Inspect

Figure 4 reports task-specific height, placement, tile and background changes. Dream-Tac/Cosmos results are 90/30 at +5 cm, 75/0 at -5 cm, 80/80 for OOD placement, 85/15 for unseen tiles and 70/25 for altered background.

Go to primary source ↓
e-efficiencyPDF p. 7, Section 4.3; p. 8, Figure 5 and captionInspect

Figure 5 concerns Peel Cucumber. Full tactile+bias training times are 80.82 versus 27.48 seconds. Inference is 60%/481 ms at 1 step, 80%/972 ms at 5, 85%/1109 ms at 10 and 85%/619 ms at 10 with cache. Training aggregation and benchmark-specific hardware are not identified in the figure/caption.

Go to primary source ↓
e-timing-contextPDF p. 5, Section 3.5; p. 1, Figure 1(c); p. 8, Figure 5Inspect

Section 3.5 separately reports H200 training 97→29 ms/iteration and ten-step inference at 5 Hz on A800. Figure 1 shows Cosmos/Dream-Tac latency 987/619 ms; Figure 5 uses 1109/619 ms for uncached/cached Dream-Tac. These are distinct comparisons with incompletely reconciled timing contexts.

Go to primary source ↓
e-flashPDF pp. 13–14, Appendix B.1, Eqs. (13)–(15)Inspect

Rank-one bias is folded into augmented query/key dot products; query content is pre-scaled by 1/sqrt(d). Padding preserves logits and fused computation avoids a dense S×S bias matrix.

Go to primary source ↓
e-cachePDF p. 14, Appendix B.2; p. 15, Figure 11Inspect

Validation-set diffusion diagnostics motivate full evaluations at the first and third steps and cache reuse elsewhere. Generic intermediate-feature caching is described, but precise cache tensors and complete implementation are not specified.

Go to primary source ↓
e-gate-diagnosticPDF pp. 7–9, Section 4.4.2; p. 8, Figure 7; p. 13, Appendix A.6/Figure 10Inspect

Five sampled Peel Cucumber training episodes contribute 874 noninitial timesteps. Figure 7 plots blue tactile change and red gate with episode 75th-percentile dotted lines. Text attributes approach oscillations to noise/disturbance and describes phase tracking without independent contact labels.

Go to primary source ↓
e-latentPDF pp. 7–8, Section 4.4.1 and Figure 6Inspect

Pretrained Wan VAE tactile embeddings are visualized with t-SNE clusters for manipulation actions; no quantitative separability or causal control test accompanies the plot.

Go to primary source ↓
e-reconstructionPDF pp. 14–15, Appendices C.1–C.2; p. 16, Figures 12–13Inspect

Cut Banana examples compare predicted and ground-truth visual frames and tactile images. Figure 13 contrasts pre-contact/cutting and highlights lower tactile regions. The source acknowledges smoothing; it reports qualitative correspondence rather than reconstruction-error metrics.

Go to primary source ↓
e-task-definitionsPDF p. 6, Section 4.1 task list; p. 11, Appendix A.2Inspect

Main-text success requires a peeled strip longer than 5 cm and cut depth exceeding 5 cm. Appendix A.2 describes clean skin removal and clear banana separation. Pick and Place is the baguette task.

Go to primary source ↓
e-limitationsPDF p. 15, Appendix DInspect

Authors identify limited task/environment coverage, limitations of frame-to-frame gating and remaining diffusion expense. Broader datasets, richer contact modeling and harder long-horizon/dexterous settings are future work.

Go to primary source ↓

8.5 Primary sources

Scroll across the image to inspect details. Press Esc to close.