TourPhysics: Bringing Physics to World Models for Exploration and Manipulation from a Single Image
1. Paper overview
In one sentence: TourPhysics keeps simulation authoritative while diffusion fills appearance, gaining persistent visual control at the cost of a declared, incomplete scene hypothesis and non-real-time generation. identityoverviewtransactiondepthmemorymotion-tablememory-evaluationgenerationlimits
| At a glance | What to know |
|---|---|
| Research problem | Source description A convincing video can change a revisited scene or ignore an intervention. TourPhysics asks how to preserve a declared physical hypothesis while expanding visible appearance. Single-image ambiguity makes scale, hidden geometry and material parameters assumptions of that hypothesis. identityoverviewlimits |
| Core mechanism | Author claim The authors extend PhysOmni with persistent action–observation transactions and explicit ownership of simulator state, geometry, generator controls and accepted appearance. identityoverviewtransaction |
| A key reported result | Prescribed camera-motion adherence: 0.0233 simulator units; 1.4301; 98.26% / 93.00%. ATE ↓; PLR → 1; T-Dir/R-Dir ↑. 49 sequences, 40 identities; ViPE recovery with one whole-sequence Umeyama similarity alignment. LingBot-Cam: 0.1031 ATE. MotionCtrl has closer PLR (0.8517). Sora's 99.99% R-Dir has one eligible segment. Lowest displayed ATE, measured in declared scene units. Direction hits test half-space agreement. Controls and temporal support differ across methods. protocolmotion-tablecamera-protocol |
| Reading caution | Source description Declared proxies and parameters do not recover unique real-world physics. Hidden geometry remains incomplete; camera/depth errors reduce correspondence. Static-only memory cannot preserve deforming-object appearance generally. Diffusion is not real-time, and uncompressed pages grow with accepted views. limits |
Core contributions
- Author claim
The authors extend PhysOmni with persistent action–observation transactions and explicit ownership of simulator state, geometry, generator controls and accepted appearance. identityoverviewtransaction
Figure 2. A fixed simulated trajectory becomes a generated observation; accepted evidence conditions later windows. Original paper, p. 5 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. Begin with the top row: masks, meshes, a background plate and calibrated depth create one persistent hypothesis from the image. The declaration on the right provides material assignments, forces and camera commands. Follow the middle row from the orange physical transition into the blue geometric record, then the purple observation model and decoded output. Geometry is already fixed when synthesis starts. The lower green route carries accepted appearance history back into conditioning, while the lower-right depth branch prepares later controls. Read these return arrows together with Equations (3)–(5): candidate memory stays private until acceptance, and neither return path rewrites simulator geometry. overviewtransactionliftingphysicslimits
What it supports. The diagram's useful distinction is ownership. Simulation determines the candidate physical history; the generator contributes appearance; acceptance advances both persistent state and visual evidence together. This allows later observations to reuse generated knowledge while retaining one declared physical trajectory. It does not make the original single-image reconstruction uniquely correct.
Where the evidence stops. The filmstrip is illustrative evidence of the intended output. It cannot establish continuous motion quality or real-world physical accuracy. Camera-only actions may still advance gravity-driven dynamics; absence of a new intervention does not imply a frozen world.
2. Motivation
2.1 The problem and the proposed response
A convincing video can change a revisited scene or ignore an intervention. TourPhysics asks how to preserve a declared physical hypothesis while expanding visible appearance. Single-image ambiguity makes scale, hidden geometry and material parameters assumptions of that hypothesis. identityoverviewlimits
2.2 What this reading follows
A camera turn and a push both change an image, but they have different causes. TourPhysics makes those causes explicit: a simulator computes the physical and camera trajectory, then a video model synthesizes what that fixed trajectory would look like. A completed window must pass a publication gate before its endpoint and appearance evidence become persistent. This reading follows the architecture through two distinct feedback channels—relative-depth completion and static appearance memory—and then separates motion adherence from physical-event fidelity. The supplied v2 paper reports strong camera tracking and smaller held-out memory gains, while leaving real-world accuracy and a complete observation-model training recipe unresolved. identityoverviewtransactiondepthmemorymotion-tablememory-evaluationgenerationlimits
3. Research context
We place the paper in the collection through its world–action interface. The catalog labels and the reading’s assessment are shown separately.
| Catalog dimension | Recorded classification |
|---|---|
| Major category | Not assigned |
| Architecture | Not assigned |
| Prediction paradigm | Not assigned |
| Quadrant | Not assigned |
This table preserves the labels recorded at reading time. The current major category is Benchmarks & simulators. View the current classification.
3.1 Evidence-based assessment
Insufficient evidence to decide
The catalog has no assigned classification to affirm or reject. Architecturally this is a modular simulator-plus-observation generator with inference-time appearance feedback. User actions are inputs, not predicted actions; neither joint future/action prediction nor inverse-dynamics action extraction is established. A One Model action-policy quadrant is therefore unsupported. overviewtransactionphysicsgeneration
This is the collection’s architectural analysis, not a new related-work survey. Benchmark comparisons and their protocols appear in Section 6.
4. Problem formulation
4.1 Inputs and outputs
| Inputs | Outputs |
|---|---|
|
|
4.2 Equations and their role
5. Method in detail
5.1 One action creates a candidate history before any pixels are sampled
The persistent state contains physical state, camera state, appearance memory and accepted tail evidence. For a new action, the simulator starts from a clone, computes the full physical/camera segment and exports a read-only geometric record. This makes the observation model's task conditional: it must render a trajectory whose timing and geometry are already chosen. The generator does not infer the action to execute. After decoding, new memory pages and tail estimates are staged privately. Acceptance publishes the terminal states and staged evidence together; rejection leaves the committed index unchanged. Appendix A.5 specifies explicit retries, with only the sampling seed allowed to change. The overlap boundary is checked in RGB, but the simulator endpoint follows the sampled index convention. This is why retry handling must restore both caches and physics rather than merely discard an unattractive video. transactiongatetiming
Figure 5. Accepted tail evidence can fill eligible control holes without changing metric geometry. Original paper, p. 11 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. Read the upper blue boxes first: the fixed geometric record produces the base generator control. The orange boxes align previously accepted tail estimates and filter them for eligibility. In the purple fusion box, the base receives weight one minus W, and historical completion receives weight W. The lower panels visualize base control, mask, aligned estimate and fused control. The white and gray mask legend marks eligible and unchanged regions. The yellow fallback panel should be read through Equation (12): when W is zero, generated control equals base control. These are normalized depth conditions, whereas simulator z-depth retains its metric projection and visibility role. depthtransactionablation-tablefallback
What it supports. This branch changes what depth guidance the diffusion model sees in under-supported static regions. It does not accumulate generated depth as physical measurement. The fail-closed tests mean missing overlap, poor agreement or a degenerate fit removes the proposed completion rather than allowing it to alter the simulator's geometric record.
Where the evidence stops. The caption identifies these depth maps as illustrative. Figure 5 shows binary mask endpoints, while Appendix A.3 permits weights in [0,1]. The graphic establishes routing; it supplies no isolated quantitative ablation of depth completion.
5.2 Two appearance feedback channels solve different problems
The original image is the permanent appearance anchor, but it cannot show every surface revealed by a tour. TourPhysics handles that gap with two bounded channels. First, accepted tail evidence proposes relative-depth completion for eligible holes in a later generator control. Metric geometry remains unchanged. Second, clean attention pages carry appearance from earlier accepted static views. Target pixels are back-projected with metric depth, projected into historical cameras and checked for visibility, static support and depth consistency. The best eligible page contributes through a bounded residual around native attention. Failure in the first channel sets the completion weight to zero; failure in the second sets the memory gate to zero. These fallback rules protect the native conditioning path. They do not solve missing physical geometry or create memory for deforming objects, which the paper explicitly excludes. liftingdepthmemorylimits
Figure 6. History is readable only after commit and can be switched off by geometry. Original paper, p. 13 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. Panel (a) separates storage from availability: a clean page becomes committed only after the enclosing observation passes the quality gate. The native cache illustration has eight reference-sink slices and four recent slices. Its key-rephasing arrow changes the temporal rotary position from source time to just before target time; the spatial tokens and stored attention-value tensor remain unchanged. Panel (b) compares native attention with a page-augmented historical read. Follow both outputs into the final residual box. Geometry contributes validity and confidence to the gate, and the bottom line states the exact invalid-correspondence fallback. The caption distinguishes stored attention values from geometric visibility. memory-diagrammemorytransactionfallbacklimits
What it supports. The mechanism retrieves appearance associated with earlier static surfaces instead of treating every old frame as universally relevant. Its bounded residual makes zero valid correspondence an exact native-attention case. Candidate pages cannot influence other frames in their own uncommitted window; the first opportunity to use them is a later accepted window.
Where the evidence stops. Exact fallback is an algebraic property and an author-reported test, not proof that every accepted correspondence is correct. Static-only eligibility excludes moving and deforming objects, and the growing page store has no demonstrated fixed-capacity solution.
5.3 Evaluate the simulation hypothesis, then isolate the memory effect
Reader analysis: the broad comparison asks whether each system can follow TourPhysics's declared simulator target through its supported interface. It does not identify real scene parameters or match every method's control information. Camera error is aligned in internal units; object and event scores depend on tracking and visibility, with especially small supports for response and recovery. The memory experiment supports a narrower causal interpretation because its treatment and control hold trajectories, generator depth, diffusion settings and seeds fixed. Its held-out identities are separate from development, but eligible scenes still require the same revisit structure. Historical error measures consistency with generated history, while known-background error is a separate simulator-based diagnostic. Both are needed: a system could otherwise preserve an earlier appearance mistake very consistently. Confidence intervals and raw paired outcomes would help assess the smaller held-out gain. protocolcamera-protocolobject-protocolinteraction-protocolmemory-evaluationlimits
5.4 Training and inference
During training
Appendix A.4 describes a finite-window diffusion model trained for geometry-conditioned synthesis. The supplied paper does not specify its exact backbone/checkpoint, training corpus, loss, optimizer, training schedule or hardware budget. Frozen weights during sampling do not establish that the entire method is training-free. generation
During inference
Controls are 640×352; decoded windows contain 81 frames at 832×480 and 16 fps, sharing one boundary and adding 80 frames. Simulator records are indexed at 60 fps using center-of-bin sampling; the committed endpoint is the sampled shared boundary. Later duplicate boundary frames are removed. generationtimingtransaction
The gate checks decoding, boundary continuity, black pixels and frozen motion. Its optional reprojection manifest was absent, and simulator RGB errors were diagnostic only. There is no automatic retry: an explicit retry preserves action, trajectory, controls and committed memory, permitting only a changed sampling seed. gate
5.5 Implementation flow
- Lift one scene hypothesis
SAM 3 masks and SAM 3D Objects meshes provide explicit proxies. LaMa forms a background plate. Ground alignment, silhouette-based camera refinement and affine monocular-depth calibration establish shared geometry. The original image remains an immutable appearance anchor. lifting
- Simulate before generating
Clone the committed physical and camera state. Genesis routes rigid bodies to RBD, continuum materials to MPM and thin constraint systems to PBD. The complete camera/physical segment and geometric record are fixed before synthesis. Gravity and ongoing dynamics can advance even without a new intervention. transactionphysics
- Separate geometry from conditioning
Metric camera-space z-depth governs projection, visibility and evaluation. Accepted tail estimates can complete only eligible, newly exposed static holes in normalized generator depth; failed overlap, affine-fit or agreement tests preserve the base control. depth
- Retrieve accepted static appearance
Back-project target pixels with metric depth and camera transforms, then test visibility, static support, confidence and depth agreement in earlier committed views. Select at most one eligible page per target location. Clean attention features affect a bounded residual; dynamic objects are excluded. memorymemory-diagram
- Publish atomically
Stage new appearance pages and tail evidence privately. Acceptance publishes these together with terminal physical/camera states. Rejection preserves the checkpoint and committed index. New pages and tail evidence can first affect a subsequent window. transactiongatememory
6. Experiments & results
TourPhysics renders persistent camera tours and object manipulations from a single image plus a declared physical scene. A simulator fixes each trajectory before diffusion generates its appearance; accepted windows publish physical endpoints and visual memory together. The strongest evidence concerns prescribed motion and modest improvements on held-out revisits, with limited physical-event support and no real-world validation.
6.1 Read the original evidence
Table 2. Camera tracking and visible-object alignment improve under heterogeneous control interfaces. Original paper, p. 16 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. Keep the two column groups separate. Camera ATE is position error after one whole-sequence similarity alignment; PLR measures recovered versus declared path length, so closeness to one matters. T-Dir and R-Dir count half-space direction hits, not exact rotation or translation recovery. The object columns compare tracked centers in generated and simulator videos, normalized by image diagonal. S@0.05 thresholds each retained object's ADE before averaging within sequence and across sequences. The footnote matters as much as the bold cells: TourPhysics retains 168 objects across 93 sequences, and each method has its own visibility support. Table 1 separately documents unequal input controls. motion-tablecamera-protocolobject-protocolprotocol
What it supports. TourPhysics reports ATE 0.0233 versus LingBot-Cam's 0.1031, with T-Dir 98.26% and R-Dir 93.00%. Object ADE is 0.0297 versus Tora-Initiator's 0.0462. However, MotionCtrl's PLR 0.8517 is closer to one than TourPhysics's 1.4301, so the result is not dominance across every camera diagnostic.
Where the evidence stops. ATE uses internal simulator units, not recovered meters. Sora's displayed 99.99% R-Dir has one eligible segment and is excluded from best highlighting. Stretched short baselines and method-specific visibility prevent a fully matched ranking.
Table 3. Physical-event adherence has strengths and conspicuous remaining failures. Original paper, p. 17 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. Read collision proxies, deformation proxies and perceptual diagnostics as separate questions. Contact error estimates event timing from tracked point-set approach; response accuracy asks whether a target moves after contact. Shape error uses normalized Procrustes alignment, area drift compares projected hull-area ratios, and recovery checks deformation amplitude plus post-event relaxation. The rightmost VideoPhy-2 scores assess semantic adherence and physical commonsense on a 1–5 scale; they do not verify the prescribed trajectory. Preserve the support superscripts when comparing rows. The footnote and Table 6 show that contact, response, shape and recovery are scored on different subsets, then reduced within and across sequences. interaction-tableinteraction-protocol
What it supports. TourPhysics has the lowest displayed contact error, 3.3304 seconds, and area drift, 0.0866. Its response accuracy is only 16.67%, below Wan and Sora at 33.33%, and its shape error 0.0704 exceeds LingBot-Cam's 0.0664. The table therefore supports partial adherence rather than a blanket physical-fidelity claim.
Where the evidence stops. TourPhysics response covers five sequences, and recovery covers eight events from two identities. Tora's near-perfect recovery has one event. Near-zero excess 2D overlap cannot certify zero 3D penetration; area drift is not mass conservation.
6.2 Results and evaluation conditions
| Task & protocol | Reported result | Comparison & interpretation |
|---|---|---|
| Prescribed camera-motion adherence 49 sequences, 40 identities; ViPE recovery with one whole-sequence Umeyama similarity alignment. | 0.0233 simulator units; 1.4301; 98.26% / 93.00%. ATE ↓; PLR → 1; T-Dir/R-Dir ↑ | LingBot-Cam: 0.1031 ATE. MotionCtrl has closer PLR (0.8517). Sora's 99.99% R-Dir has one eligible segment. Lowest displayed ATE, measured in declared scene units. Direction hits test half-space agreement. Controls and temporal support differ across methods. protocolmotion-tablecamera-protocol |
| Prescribed object-motion adherence CoTracker3, frozen masks, 8 fps, 682×384; 168 objects/93 sequences/49 identities, requiring eight jointly visible frames and 15% joint visibility. | 0.0297 / 0.0524; 81.18%. ADE / endpoint error ↓; S@0.05 ↑ | Tora-Initiator: 0.0462 / 0.0865; 70.05%, with oracle initiator tracks and different valid support. Errors normalize by image diagonal; object statistics are averaged within sequence, then equally across sequences. S@0.05 is not robot-task success. protocolmotion-tableobject-protocol |
| Simulator-defined contact and deformation Paired image-space tracking; contact: 14 sequences/6 identities; response: 5/3; shape/area: 30 records/25 sequences/13 identities; recovery: 8 events/7 sequences/2 identities. | 3.3304 s; 16.67%; 0.0704 / 0.0866; 64.29%. Contact error ↓; response ↑; shape/area errors ↓; recovery ↑ | minWM-Wan contact: 4.5639 s. Wan/Sora response: 33.33%. LingBot-Cam shape: 0.0664. Tora recovery >99.99% covers one event. Mixed results on method-specific supports. Low projected overlap does not establish collision-free 3D physics; VideoPhy-2 scores are auxiliary perceptual diagnostics. interaction-tableinteraction-protocol |
| Persistent-cache brightness drift 72 paired videos from 27 identities; fixed observation backbone and generator. | Persistent sink: 11.17 / 7.74. End-to-start / tail-to-head drift ↓, 0–255 intensity levels | Ordinary cache: 26.33 / 20.38. Reduced drift of mean decoded channel values, not a percentage improvement or a geometric-consistency metric. ablation-tablecache-protocol |
| Held-out appearance-memory revisits 14 identity-disjoint dynamic scenes; six windows, 481 frames, 51 planned target latents, revisit chunks 3–5; matched trajectories, depth, checkpoints, prompts and seeds. | −2.39% / −0.90%. Historical / known-background error change ↓ | Memory disabled is the paired control; four development identities yield −8.25% / −2.99%. History uses mean paired percentage changes; held-out background uses the percentage change of identity-macro MAE. Development selection was biased toward informative revisits. No uncertainty intervals are supplied. ablation-tablememory-evaluation |
| Input-scene and camera-class diagnostics Section 5.4 and Appendix B.4 report TourPhysics-only diagnostics. | 0.9743 / 0.9791 / 0.8921. Subject consistency / background consistency / camera-motion class accuracy ↑ | No corresponding baseline values are supplied. The cited appendix repeats these scores without sufficient estimator and denominator details to reconstruct their protocol. appearance-diagnostics |
6.3 Ablations and diagnostic examples
Read component removals and qualitative examples within their stated evaluation conditions.
Table 4. Cache stability and geometrically routed history are tested as distinct interventions. Original paper, p. 18 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. First compare only the upper two rows: they keep the observation backbone and generator fixed while changing the temporal cache. End-to-start and tail-to-head drift measure differences in mean decoded 8-bit channel values, with no percentage normalization. Then move to the lower two rows, where memory-on is compared with memory-off under matched trajectories, generator depth, prompts, checkpoints and seeds. Negative percentages mean lower error. The history metric follows static depth-consistent correspondence; known-region error uses the simulator-consistent background mask. Appendix B.4 specifies different aggregation for the held-out background column, so its percentage must not be interpreted as the mean of per-scene percentages. ablation-tablecache-protocolmemory-evaluation
What it supports. The persistent sink cache lowers the two brightness-drift diagnostics from 26.33/20.38 to 11.17/7.74. Appearance memory reduces historical error by 8.25% on development scenes and 2.39% on held-out identities; corresponding known-region changes are −2.99% and −0.90%. The held-out effect is favorable but smaller.
Where the evidence stops. Development selection favored interpretable revisits and used prior screening information. The held-out set has 14 dynamic identities and excludes camera-only scenes. These means lack uncertainty intervals; neither brightness stability nor historical consistency alone establishes accurate physical dynamics.
7. Analysis & limitations
7.1 What the evidence leaves open
Declared proxies and parameters do not recover unique real-world physics. Hidden geometry remains incomplete; camera/depth errors reduce correspondence. Static-only memory cannot preserve deforming-object appearance generally. Diffusion is not real-time, and uncompressed pages grow with accepted views. limits
Unequal conditioning, stretched short outputs and method-specific visibility filtering limit cross-system rankings. Physical-event supports are especially small. Within-model memory ablations isolate the mechanism more directly than the baseline comparison. protocolobject-protocolinteraction-protocolmemory-evaluation
7.2 Questions for discussion
- How would rankings change under common visibility support and matched control information?
- Can bounded object-local memory preserve deforming surfaces without corrupting geometry?
8. Reproducibility audit
8.1 Requirements and known gaps
Reconstruction requires the named scene front end, Genesis, observation weights, frozen declarations, pose exports, memory plans and paired evaluation records. Defaults include a 4 ms solver step, 10 rigid-only or 16 particle substeps; engine versions, compute costs and several memory thresholds remain unspecified. liftingphysicsgenerationmemorymemory-evaluation
Proposed checks: force a rejected window and verify unchanged physical/cache/history state; then compare memory-off, correct correspondence and deliberately invalid correspondence with matched seeds and controls. The illustrated edition specifies controls and falsifiable outcomes. gatememorymemory-evaluationfallback
8.2 Proposed reproduction checks
The following checks are proposals motivated by the paper. They have not been run as part of this reading.
Check 1: Reject a window and test the complete rollback boundary
Reader-proposed check, not performed: begin two runs from the same committed checkpoint and fixed action. In one run, force a media-contract rejection after candidate memory and tail evidence have been staged. Compare physical/camera state, native cache, appearance pages, accepted evidence and committed index with the untouched control checkpoint. Then explicitly retry with the control's sampling seed. Verify identical simulator records and controls, and compare the accepted continuation with the direct control run. Include the center-of-bin shared endpoint in the comparison. Any extra physical advance, surviving staged page, altered control or shifted boundary falsifies the claimed transaction isolation. transactiongatetimingmemory
Check 2: Separate useful correspondence from the mere presence of history
Reader-proposed check, not performed: on the identity-disjoint six-window revisit set, compare memory-off, correctly routed memory and memory with all correspondence masks invalidated, holding trajectories, depth, native cache, checkpoints and seeds fixed. At the attention output, the invalid-mask arm must recover the native path; compare decoded outputs under deterministic execution as a secondary check. Add a stress arm with deliberately mismatched source-view camera/depth records while retaining eligibility tests. Record gate coverage, historical error on a frozen valid evaluation mask and known-background MAE, reporting each identity separately. No fallback equivalence, or historical gains accompanied by background degradation, would weaken the proposed mechanism. memorymemory-evaluationfallback
8.3 Reading coverage
Visual audit: Actually inspected the title/byline/version page, all eight original figures, all six tables, all method/evaluation pages and Appendices A–C as rendered PDF pages. Read the entire supplied text, including reference-only pages 20–25. Inspected all six final original crops; table footnotes and graphic legends are retained. Checked Figure 2's feedback ordering against Eqs. (1)–(5), Figure 5's mask/fallback against Eq. (12) and Appendix A.3, and Figure 6's temporal rephasing and residual against its caption and Eqs. (14)–(19). The inspected claim-relevant arrows and gates agree with the formulations; Figure 5's binary illustration is distinguished from the permitted continuous weight. No external supplement, full video, code or earlier paper version was visually reviewed.
PDF pages inspected for this edition: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36. Appendix coverage: reviewed.
Original figures and tables remain the work of the source’s authors. Extractions preserve their scientific content; any HTML wrapper layout is disclosed with each figure. The surrounding reading notes are our own.
Text reading scope & known omissions
- Title, abstract and Sections 1–2 (including 2.1–2.4)
- Section 3.1–3.3: formulation and committed state
- Section 4.1–4.6: complete method
- Section 5.1–5.5: all experiments
- Section 6 and References
- Appendix A.1–A.6: implementation, simulation, controls, generation, gate and memory
- Appendix B.1–B.4: all evaluation protocols
- Appendix C.1–C.2: ablations and limitations
Outside the original text pass
- Text extraction does not reconstruct figure images; inspect the retained PDF for figures and equation/table layout.
- Reviewed arXiv:2609.04911v2, dated 7 September 2026. Title and all seven authors match the catalog; its 4 September submission date precedes this revision. No earlier version was supplied, so inter-version scientific changes were not compared.
- Text extraction does not reconstruct figure images; this was addressed by inspecting the supplied PDF, including every figure and table. Reference-only pages 20–25 were read as text, without a separate visual pass.
- Separate supplemental material availability has not been fully verified.
- External references, code, checkpoints and full videos were not inspected; experiments were not reproduced.
The visual audit above records the subsequent illustrated pass.
8.4 Traceable evidence
identityPDF p. 1, title/byline, arXiv version stamp and abstract
Exact title; Xin Zhang, Yabo Chen, Zixuan Duan, Haibin Huang, Chi Zhang, Feng Xu, Xuelong Li; v2, 7 Sep 2026. The abstract identifies the PhysOmni extension.
Go to primary source ↓overviewPDF pp. 2–5, Sections 1 and 3, Figures 1–2 and captions
Single-image declarative initialization; distinct physical, observation and epistemic branches; persistent state and immutable reference.
Go to primary source ↓liftingPDF pp. 7–8, Section 4.2; p. 26, Appendix A.1, Eqs. (A1)–(A5)
SAM 3, SAM 3D Objects, LaMa, ground alignment, camera refinement and affine depth calibration initialize a fixed hypothesis.
Go to primary source ↓physicsPDF p. 8, Section 4.3; pp. 27–28, Appendix A.2
Genesis solver routing, optional initialization-only Qwen2.5-VL-72B-Instruct, coupled contacts and simulation defaults; dynamics can continue without intervention.
Go to primary source ↓transactionPDF pp. 4–7, Sections 3.1–3.3, Eqs. (1)–(6); p. 29, Eqs. (A15)–(A18)
Committed state, cloned trajectory, conditional generation, staged updates, acceptance/rejection and boundary deduplication.
Go to primary source ↓depthPDF pp. 9–11, Section 4.4, Eqs. (11)–(12), Figure 5; p. 28, Appendix A.3
Metric geometry and normalized control have distinct writers; eligibility-gated tail completion preserves base control outside valid holes.
Go to primary source ↓memoryPDF pp. 11–12, Section 4.5, Eqs. (14)–(19); p. 30, Appendix A.6
Static geometric correspondence, confidence, one-page routing, bounded residual, clean features and delayed publication.
Go to primary source ↓memory-diagramPDF p. 13, Figure 6 and caption
Clean page commits after quality acceptance; native W12/S8 cache; temporal key rephasing leaves spatial tokens and attention values unchanged; invalid reads fall back.
Go to primary source ↓generationPDF p. 11, Section 4.5; pp. 28–29, Appendix A.4
Geometry-conditioned finite-window diffusion; fixed sampling weights/schedule; output/control sizes. No complete training or hardware recipe is provided.
Go to primary source ↓timingPDF p. 9, Eq. (10); p. 28, Eq. (A11) and boundary convention
60-fps record, 16-fps output, 81-frame windows with 80-frame stride; sampled endpoint defines the committed boundary.
Go to primary source ↓gatePDF p. 13, Section 4.6, Eqs. (20)–(21); pp. 29–30, Appendix A.5
Media, overlap, boundary-motion, black/frozen tests; absent optional manifest; diagnostics excluded from acceptance; explicit retries restore committed state.
Go to primary source ↓protocolPDF pp. 14–15, Section 5.1, Table 1 and note; p. 31, Appendix B.1
108 sequences/61 identities; heterogeneous generation controls and durations, oracle Tora tracks, and time-stretched Wan/MotionCtrl outputs.
Go to primary source ↓motion-tablePDF p. 16, Table 2, all camera/object columns and footnote
Reported motion values, method-specific visibility, TourPhysics support and Sora single-segment rotation caveat.
Go to primary source ↓camera-protocolPDF p. 31, Appendix B.2, Eqs. (B1)–(B4); p. 32, Table 5
Whole-sequence alignment, internal-unit ATE, path-length ratio, half-space direction test and support counts.
Go to primary source ↓object-protocolPDF pp. 31–32, Appendix B.2, Eqs. (B5)–(B9), Table 5
Frozen-mask tracking, normalized center error, visibility thresholds and within-sequence then equal-sequence aggregation.
Go to primary source ↓interaction-tablePDF p. 17, Table 3, all columns and footnote
Contact/deformation values, small supports, singleton Tora recovery and separate VideoPhy-2 diagnostics.
Go to primary source ↓interaction-protocolPDF pp. 32–34, Appendix B.3, Eqs. (B10)–(B13), Table 6
Contact proxy and missing-approach penalty; response, projected overlap, normalized deformation/recovery definitions and effective supports.
Go to primary source ↓ablation-tablePDF pp. 18–19, Table 4 and Section 5.5
Cache and memory comparisons, paired controls and all reported drift/error changes.
Go to primary source ↓cache-protocolPDF p. 34, Appendix B.4, Eqs. (B14)–(B16)
Mean decoded-channel drift, 0–255 units, equal-weight means over 72 paired sequences/27 identities.
Go to primary source ↓memory-evaluationPDF pp. 34–35, Appendix B.4, Eqs. (B17)–(B18) and Memory subsets
Historical Charbonnier error, background MAE, distinct percentage aggregations, four selected development and 14 held-out dynamic identities.
Go to primary source ↓appearance-diagnosticsPDF p. 17, Section 5.4; p. 34, opening of Appendix B.4
Three appearance/camera-class scores are repeated without complete estimator or effective-support specifications.
Go to primary source ↓fallbackPDF p. 35, Appendix C.1, final paragraph
Authors report disabling valid historical correspondences recovers the native observation path.
Go to primary source ↓limitsPDF p. 36, Appendix C.2
Hypothesis-dependent physics, incomplete geometry, dynamic-memory exclusion, non-real-time diffusion and growing page storage.
Go to primary source ↓8.5 Primary sources
TourPhysics: Bringing Physics to World Models for Exploration and Manipulation from a Single Image ↗
PDF · 17,377 extracted words
Source fingerprint
0c7d3dc78958996919e70b7dfdc32b5213c792642a23205edb5d1715c27b8c01