From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence
1. Paper overview
In one sentence: The roadmap proposes reusable physical reasoning through explicit consequence, intent and execution contracts, while leaving transfer benefits and the final model architecture to empirical testing. e-contracte-ownershipe-agendae-conclusion
| At a glance | What to know |
|---|---|
| Research problem | Source description Physical-intelligence progress does not automatically accumulate when model outputs, datasets and controllers encode incompatible assumptions. The authors distinguish stronger model capacity from capability reuse and system-level accumulation. Their three coupled gaps concern model roles and representations, objectives and standardization, and execution ecosystems. A forecast needs a declared downstream decision; a success score needs an embodiment and controller context; a reusable interaction needs the translations and failures that produced it. e-gapse-contributions |
| Core mechanism | Source description The survey separates system responsibility from the representation crossing the prediction-to-control interface. Explicit images, geometry or plans, predictive latent features, and latent transition/action codes are overlapping interface families. Their abstraction level does not establish whether the complete system is modular or jointly trained. e-role-taxonomye-control-interfaces |
| Reading caution | Source description The paper provides conceptual synthesis and prospective tests, without a new quantitative comparison, experimental ablation or deployed embodied-brain result. Its conclusion explicitly treats reuse, localized adaptation and safe learning progress as hypotheses. Dataset sizes cannot establish these benefits. e-contributionse-agendae-conclusion |
Core contributions
- Source description
The survey separates system responsibility from the representation crossing the prediction-to-control interface. Explicit images, geometry or plans, predictive latent features, and latent transition/action codes are overlapping interface families. Their abstraction level does not establish whether the complete system is modular or jointly trained. e-role-taxonomye-control-interfaces
- Source description
The proposed WAM contract requires a decision-relevant consequence with horizon, reference frame, uncertainty and validity conditions. The embodied-brain target additionally integrates context and selects intent. Neither RGB generation nor an action vector alone defines that predictive contract. e-contract
- Source description
Embodiment, Task and Trace Cards link local hardware constraints to goals and decision histories. Layered evaluation and admission gates are intended to make these histories useful for learning and diagnosis across systems while retaining original artifacts and embodiment details. e-cardse-learninge-evaluation-layers
Figure 2. Prediction, intent and execution expose different information at their boundaries. Original paper, p. 5 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. Start with the upper panel: physical context and a candidate intervention enter the WAM prototype, which exposes a consequence together with horizon, frame, uncertainty and validity conditions. Then follow the lower panel's solid blue route. Context enters the embodied brain; its intent reaches the harness, which calls tools and controllers. The observed outcome returns through the Trace Card to evaluation and admission. The lower green dashed arrows carry admitted experience back to the brain and harness. The gray link identifies a predictive capability under study, rather than specifying a mandatory separately deployed WAM. Read these roles together with Table 6 and Section 4.2. e-stack-figuree-contracte-ownershipe-conclusion
What it supports. The useful unit of reuse is an interpretable interface: what is predicted, what is intended and what was executed must remain distinguishable. This separation would let an evaluator attribute an outcome to model reasoning, translation or execution, but the diagram describes the proposed stack rather than demonstrating that attribution experimentally.
Where the evidence stops. The upper WAM-to-consequence arrow is green dashed, although the legend labels that style learning feedback. The caption and Section 4.2 describe a prediction output. This unresolved style inconsistency is retained; it should not be read as an extra training step.
2. Motivation
2.1 The problem and the proposed response
Physical-intelligence progress does not automatically accumulate when model outputs, datasets and controllers encode incompatible assumptions. The authors distinguish stronger model capacity from capability reuse and system-level accumulation. Their three coupled gaps concern model roles and representations, objectives and standardization, and execution ecosystems. A forecast needs a declared downstream decision; a success score needs an embodiment and controller context; a reusable interaction needs the translations and failures that produced it. e-gapse-contributions
2.2 What this reading follows
A robot can choose a useful motion without exposing what it expects that motion to change. This survey asks how such expectations could become reusable across controllers, bodies and tasks. Its answer links a predictive WAM interface to an embodied-brain intent interface, a physical harness and replayable experience records. The important distinction is between estimating a consequence, requesting a change and actually realizing that change. Read the diagrams as proposed responsibilities and the tables as a map of evidence requirements. They outline experiments worth building; they do not report a trained embodied brain or demonstrate that standardized interfaces already deliver general physical intelligence. e-contracte-ownershipe-agendae-conclusion
3. Research context
We place the paper in the collection through its world–action interface. The catalog labels and the reading’s assessment are shown separately.
| Catalog dimension | Recorded classification |
|---|---|
| Major category | Foundational work |
| Architecture | Not applicable |
| Prediction paradigm | Not applicable |
| Quadrant | Not applicable |
This table preserves the labels recorded at reading time. The current major category is Related resources. View the current classification.
3.1 Evidence-based assessment
Supports the recorded classification
The recorded foundational survey classification is supported. Architecture, prediction paradigm and quadrant are not applicable to this roadmap as a whole: it proposes functional ownership and permits modular or jointly trained realizations. Discussion of joint video/action models does not make this work itself a One Model architecture, and discussion of latent or inverse-dynamics interfaces does not assign it one action-prediction mechanism. The embodied brain is a long-term target, not a specified trained network. e-contributionse-role-taxonomye-control-interfacese-ownershipe-contract
This is the collection’s architectural analysis, not a new related-work survey. Benchmark comparisons and their protocols appear in Section 6.
4. Problem formulation
4.1 Inputs and outputs
| Inputs | Outputs |
|---|---|
|
|
5. Method in detail
5.1 Ask what a prediction changes in the decision
The paper's central distinction begins before execution. A WAM query couples the current physical context to a candidate intervention and returns a consequence with declared horizon, frame, uncertainty and validity. A brain then compares interventions and communicates the change it wants realized. These are different interfaces even if one network implements both. The survey's UniPi, VPP and LAPA examples show why representation alone cannot settle the question: explicit visual transitions, predictive features and latent action codes reach control through different mappings. Reader interpretation: the discriminating test is whether changing the consequence information changes intervention selection or verification under controlled execution conditions. Photorealistic video or accurate action imitation can be useful, but neither alone demonstrates that a consequence estimate supported the decision. e-contracte-control-interfacese-evaluation-boundarye-agenda
Figure 4. System roles and predictive representations are separate classification questions. Original paper, p. 7 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. Read the upper row as responsibilities, not a timeline. An action model emphasizes commands, a VLA emphasizes semantic action generation, a WAM emphasizes intervention consequences, and the brain target compares interventions and forms intent. The lower row asks what predictive information reaches control: an explicit forecast, predictive features or a compact transition/action representation. To make the distinction concrete, Table 1 describes UniPi translating visual transitions through inverse dynamics, VPP conditioning inverse dynamics on predictive features, and LAPA grounding latent codes with robot-action labels. Those mechanisms can occupy different interface categories without dictating whether their complete architectures share parameters. e-role-taxonomye-control-interfacese-contracte-ownership
What it supports. A latent action is not automatically an executable robot command, and a future image is not automatically a control policy. The taxonomy directs attention to the mapping from a predictive representation into a controller. It also explains why this survey itself cannot be assigned a single joint-model or inverse-dynamics quadrant.
Where the evidence stops. These categories overlap. The graphic's compact labels are not exhaustive definitions: its world-model box says future image, while the text permits broader representations. Table 1 also qualifies pi 0.7's placement by its optional visual subgoals, not its overall architecture.
5.2 Carry a peg-driving intention through the harness
Figure 16 gives a concrete conceptual example: drive a peg. The harness decomposes that request into acquiring a hammer, aligning hammer and peg, and executing a constrained impact. Arm motion drives the gripper, the gripper holds the external implement, and the implement acts on the peg and fixture. Figure 14 supplies the missing operational responsibilities around that chain: entity/frame alignment, capability resolution, precondition checks, synchronization and monitoring. The capability contract must expose frames, limits and failure modes so the intended effect remains interpretable at each boundary. This is the source's illustration, not a reported robot trial. Reader interpretation: replacing a controller while preserving the intent is a stronger test of reusable reasoning than success with one permanently fixed controller, especially if adapter effort and residual loss are recorded. e-pege-harnesse-harness-figuree-agenda
Figure 14. The harness makes an intended change meaningful to a particular embodiment. Original paper, p. 21 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. Follow the upper path from left to right. Intent must first be grounded, then aligned to entities and frames, resolved against a capability registry, checked against preconditions and synchronized with controllers. Runtime memory and the registry support different responsibilities beneath that path. On the right, tools and the environment produce outcomes that return to monitoring and recovery. The blue curve then leads to the Trace Card recorder. The green dashed return is labeled as admitted traces improving adapters and routing; it is a learning path, not the immediate control response. Section 4.3 supplies the capability-contract fields that make these boxes operationally interpretable. e-harnesse-harness-figuree-learninge-agenda
What it supports. The proposed harness is more than a tool-call dispatcher: it carries spatial, temporal and embodiment assumptions across the model-to-controller boundary. A useful test therefore measures whether substitutions preserve intent and expose failures, alongside task success. The paper also asks verifiers to reject infeasible requests without excessive false vetoes.
Where the evidence stops. The recorder-to-grounding shortcut does not display a separate admission box. Its label and Section 4.5 require evaluation before experience returns to learning. No adapter implementation, verifier threshold or measured benefit from this harness is supplied.
5.3 Turn a failed interaction into a targeted learning signal
A failed task does not identify the component that should be updated. The paper notes that valid intent can accompany controller failure, while invalid intent may appear successful in a forgiving task. Embodiment and Task Cards establish the operating conditions; a Trace Card preserves the request, translations, selected capabilities, predictions, verifier states and observed outcome. Table 8 then separates possible learning uses instead of prescribing one universal loss: predictive reasoning, interface alignment, verifier learning and control adaptation have different signals and gates. Table 9 supplies corresponding diagnostics. Reader interpretation: learning admission should depend on whether a trace supports the intended update, not merely whether its final label is success. This motivates keeping provenance, corrections and quality flags and testing regression before promoting an update; the source leaves the implementation and thresholds open. e-feedback-limitse-cardse-learninge-evaluation-layerse-agenda
Table 7. Three records connect what the robot is, what the task requires and what actually happened. Original paper, p. 23 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. Begin with the Embodiment Card: frame conventions, calibration, control rates and limits define how to interpret an action. Move to the Task Card for the allowed observations, tools, success criteria and meaningful failure labels. Then read the longer Trace Card row as a chain linking model requests to their realization. It includes representation schema and frame, candidate interventions, selected tools, adapters, predictions, verifier outputs, controller states and outcomes. A success label alone discards most of that chain. The scaling-role column describes intended uses of the records, while Section 4.4 stresses that existing raw artifacts and storage formats can be retained. e-cardse-feedback-limitse-learninge-evaluation-layers
What it supports. A reusable training example needs the context that gives its outcome meaning. These proposed records would help distinguish a bad intention from a correct intention mistranslated by an adapter or unsuccessfully executed by a controller. That distinction is central to choosing which component should receive a learning signal.
Where the evidence stops. The table gives semantic minimum fields, not a versioned executable serialization or proof of deterministic replay. Real interactions remain partially observed, and the paper provides no measured trace-completeness threshold or demonstrated cross-dataset replay result.
5.4 Training and inference
During training
Table 8 proposes complementary objectives: physical representation learning, predictive reasoning, brain–harness alignment, tool-chain supervision, verifier alignment, harness-in-the-loop reinforcement learning, and distillation/adaptation. Each pairs a supervision source with an updated component and a required evaluation gate. These are not a fixed chronological training schedule. The paper specifies no optimizer, loss equation, frozen-module recipe, hardware budget or trained checkpoint for a new system. e-learninge-ownership
The reviewed data landscape separates broad human video from embodiment-grounded robot trajectories. Action-free video can retain semantic labels while lacking target-robot commands. Table 2 reports 220,847 Something-Something V2 clips and DROID's 76,000 trajectories, 350 hours, 564 scenes and 86 tasks. These characterize cited resources, not a shared training set or benchmark score. The autonomous-interaction row leaves scale unspecified and claims no standalone public dataset. e-data-rolese-data-scale
During inference
At runtime, intent formation precedes harness grounding and execution. The harness may invoke prediction or verification models as needed, then resolve a tool chain and execute an approved action. Observation of the outcome closes the physical loop. Actuator commands and high-frequency stabilization remain downstream responsibilities; learned prediction is not evidence that an action was executed. e-ownershipe-harnesse-learning
An evaluator separately decides whether recorded experience is suitable for learning. Table 9 distinguishes predictive reasoning, grounding, policy/control, orchestration, safety, traceability and self-improvement. Some measurements exist in cited work; others are proposed. Regression and safety gates precede update promotion, with trace quality and rollback criteria remaining essential design choices. e-evaluation-layerse-learninge-agenda
5.5 Implementation flow
- Predict consequences before committing to an intervention
The proposed predictive interface can expose future observations, geometric state, latent transitions, contact changes, affordances or task progress. Its value must be established by improved intervention comparison or verification. For context, Table 1 describes UniPi coupling visual trajectories to inverse dynamics, VPP conditioning inverse dynamics on predictive features, and LAPA grounding latent transition codes with robot-action labels. These are survey summaries of distinct methods, not components of a new implementation. e-contracte-control-interfaces
- Express intent with inspectable semantics
The brain chooses an intended physical change and communicates entities, frames, constraints, uncertainty and relevant completion or recovery conditions. The encoding may be language, code, structured geometry, learned tokens or a hybrid. A WAM can contribute predictive capability without remaining a mandatory separate module. Functional separation permits joint training when intermediate transformations remain observable. e-contracte-ownership
- Ground and execute through capabilities
The harness aligns entities and frames, resolves capabilities, checks preconditions, synchronizes controllers, monitors outcomes and handles recovery. A capability contract declares inputs, coordinates, preconditions, controllable variables, outputs, timing and failures. A tool is a capability endpoint; a tool model supplies specialized inference or control. The paper's peg-driving illustration decomposes intent into acquiring a hammer, aligning it with the peg and executing a constrained impact; it reports no execution trial. e-harnesse-harness-figuree-peg
- Preserve the decision context for reuse
The Embodiment Card records morphology, sensors, calibration, frames, rates, limits and controller assumptions. The Task Card records goals, allowed observations/tools, constraints and failure labels. The Trace Card links observations and requests to candidate interventions, representation versions, adapters, tool chains, verifier/controller states, outcomes and corrections. These are semantic requirements that can accompany existing storage formats, not a released universal schema. e-cards
6. Experiments & results
This survey proposes explicit contracts linking physical prediction, model intent, execution and reusable experience. A World Action Model estimates intervention consequences; an embodied brain forms an intended change; a physical harness translates and verifies execution. The contribution is a research roadmap with testable responsibilities, rather than a trained embodied-brain system or demonstrated performance gain. Its promised reuse remains a hypothesis (e-contract, e-ownership, e-conclusion).
This is a survey and prospective roadmap, with no new empirical results table or experimental ablation. Six original visuals are retained because the source has ample conceptual material. Table 2 supplies a quantitative resource inventory, not a benchmark comparison; Table 9 supplies a diagnostic evaluation framework, not measured ablation outcomes. The remaining visuals explain roles, execution and experience records. Accordingly, the base results array is empty and no featured empirical result is selected. Proposed checks test the paper's hypotheses without claiming they were run. e-contributionse-data-scalee-evaluation-layerse-agendae-conclusion
6.1 Read the original evidence
Table 2. Resource counts describe different kinds of evidence rather than one interchangeable training pool. Original paper, p. 12 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. Read each row horizontally before comparing its size with another resource. The supervision column tells you whether observations include target-robot commands. The scope column counts different units, such as clips, hours, trajectories or tasks. The use column records the survey's account of how a resource supports prediction, grounding or evaluation. Something-Something V2 has semantic labels but no robot commands; DROID provides synchronized real-robot observations and actions. SIMPLER is explicitly separated as an evaluation framework. Finally, check availability: the autonomous-interaction row leaves scale unspecified and does not claim a standalone public dataset. Neither missing entry should be filled from another release. e-data-scalee-data-rolese-evaluation-boundary
What it supports. The survey reports 220,847 Something-Something V2 clips and DROID's 76,000 trajectories, 350 hours, 564 scenes and 86 tasks. These quantities illustrate diversity in supervision and scope. They do not measure the performance of the proposed roadmap, and more video does not itself resolve embodiment-specific action grounding.
Where the evidence stops. This is a quantitative resource inventory, not an experimental results table. Counts and availability are the survey's reports of cited sources; those external artifacts were not independently inspected. Different units, licenses, schemas and action spaces prevent pooling the rows into one comparable total.
6.2 Results and evaluation conditions
No quantitative results are included in this reading.
6.3 Ablations and diagnostic examples
Read component removals and qualitative examples within their stated evaluation conditions.
Table 9. A diagnostic framework separates existing evidence from prospective roadmap tests. Original paper, p. 24 ↗
Excerpt from the authors’ paper; cropped without altering the figure or table.
How to read it. Read the Status column before interpreting the measurements. Policy and control is marked Existing. Predictive reasoning and verification/safety are marked Both, mixing measurements used in cited literature with proposed extensions. Brain–harness grounding, tool orchestration, harness/traceability and self-improvement are Proposed. Each row then asks a different question: whether prediction captures relevant change, intent survives grounding, control succeeds, tools compose correctly, hazards are rejected, failures are attributable or updates retain safety. The table's caption defines these statuses; they describe where measurement ideas come from, not which tests this paper performed. Table 8 complements this view by associating learning objectives with required gates. e-evaluation-layerse-learninge-evaluation-boundarye-agenda
What it supports. End-task success cannot reveal which responsibility improved. The proposed diagnostics create places to look: frame consistency for grounding, false vetoes and missed hazards for verification, and regression or rollback triggers for learning. This is a useful design for future ablations, but the table contains no ablation effect sizes or new success rates.
Where the evidence stops. Existing means used in cited work, not reproduced here; Both does not mean every listed metric is established. The source gives measurement categories without a complete common protocol, numerical acceptance thresholds or evidence that a single aggregate score would be valid.
7. Analysis & limitations
7.1 What the evidence leaves open
The paper provides conceptual synthesis and prospective tests, without a new quantitative comparison, experimental ablation or deployed embodied-brain result. Its conclusion explicitly treats reuse, localized adaptation and safe learning progress as hypotheses. Dataset sizes cannot establish these benefits. e-contributionse-agendae-conclusion
Predictive fidelity, action grounding, closed-loop success and executability answer different questions. Table 3 does not support treating generated-video quality as robot success, and Table 1 warns that robots, tasks and protocols differ. A pooled ranking would exceed this survey's evidence. e-evaluation-boundarye-control-interfaces
The architecture and representation remain open. A shared world-centric 3D/4D representation is a research objective; modularity does not guarantee transfer. The central empirical burden is to show that intent survives adapter, tool and embodiment changes while reporting adaptation cost and residual loss. e-contracte-agendae-conclusion
Figure 2 uses a green dashed WAM-to-consequence arrow although its legend assigns that style to learning feedback. Its caption and Section 4.2 describe a prediction output. The illustrated edition preserves this inconsistency and follows the explicit textual contract rather than interpreting that arrow as a training operation. e-stack-figuree-contract
7.2 Questions for discussion
- Which consequence variables change intervention selection after controller quality and compute are controlled? (e-agenda)
- How much adaptation remains local when the same intent must work with a substituted controller or tool? (e-agenda)
- Which trace omissions prevent distinguishing an incorrect intent from a grounding or control failure? (e-cards, e-evaluation-layers)
8. Reproducibility audit
8.1 Requirements and known gaps
A meaningful implementation would first choose one embodiment, task, intent encoding and capability registry, then record the three Cards alongside raw observations. It would need explicit frame transformations, calibration, controller limits, verifier decisions and synchronized outcome logs. The proposed fields guide implementation, but the source supplies no complete serialized schema, executable benchmark protocol or numeric admission thresholds. e-harnesse-cardse-learninge-agenda
Reader-proposed checks should isolate the roadmap's claims: hold the task/controller fixed while removing or shuffling consequence predictions, then test controller substitution while retaining the brain request. Log prediction utility, grounding errors, adapter cost, false vetoes, missed hazards and task outcomes separately. These comparisons would test decision relevance and local adaptation; neither was performed in this reading. e-harness-figuree-evaluation-layerse-agenda
8.2 Proposed reproduction checks
The following checks are proposals motivated by the paper. They have not been run as part of this reading.
Check 1: Does consequence information improve intervention selection?
Reader-proposed experiment, not performed: choose one contact-sensitive simulated manipulation task and fix the controller, candidate interventions, observations, evaluation episodes and decision budget. Compare a consequence-aware selector with the same selector deprived of predictions and with predictions shuffled between candidates. Log declared horizon/frame, predicted contact or state change, selected intervention, verifier decision and executed outcome. Report prediction accuracy, unsafe selections, false vetoes, task success and latency separately, with uncertainty across repeated episodes. If correct consequences do not outperform both controls, or apparent gains disappear when latency is matched, the claimed decision relevance is unsupported in that setting. This tests the roadmap's first milestone without requiring a complete embodied brain. e-contracte-evaluation-layerse-agenda
Check 2: Can a controller be replaced without relearning the brain?
Reader-proposed experiment, not performed: retain the same brain and task-level requests while substituting a controller with a declared frame convention or action interface. Compare the original stack, an unadapted substitution, and a substitution using only a locally trained or calibrated harness adapter. Hold observations, allowed capabilities, task split and adaptation budget fixed. Record both Embodiment Cards and complete Trace Cards; compare intent/frame preservation, task outcomes, adapter effort, residual performance loss and failure attribution. Inject known frame-translation errors in simulation to test whether logs distinguish grounding failures from control failures. If recovery requires brain retraining or errors remain unlocalizable despite complete records, this implementation does not support the roadmap's local-adaptation hypothesis. e-harnesse-harness-figuree-cardse-evaluation-layerse-agenda
8.3 Reading coverage
Visual audit: The title/author/version page, all 16 figures and all nine tables were visually inspected. The declared pages include every source location supporting retained method, data-scale, evaluation and proposed-reproduction details, including the uncropped peg example and post-training Table 8. All six final original-PDF crops were individually viewed; table headings, diagram labels and legends are intact. Figure 2's prediction-arrow/legend inconsistency is disclosed without alteration; Figure 14's abbreviated feedback path is interpreted using its admitted-trace label and Section 4.5. Pages 2, 6 and 26–31 were read in the complete text pass but not visually inspected. No appendix was present; separate supplements and cited external artifacts were not supplied for verification.
PDF pages inspected for this edition: 1, 3, 4, 5, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25. Appendix coverage: not present.
Original figures and tables remain the work of the source’s authors. Extractions preserve their scientific content; any HTML wrapper layout is disclosed with each figure. The surrounding reading notes are our own.
Text reading scope & known omissions
- Title, author block, abstract and arXiv version stamp (p. 1)
- 1 Introduction (pp. 1–4)
- 2 Evolution of Embodied Models Toward WAMs, including 2.1–2.8 (pp. 4–13)
- 3 Current Barriers on the Path to an Embodied Brain, including 3.1–3.5 (pp. 13–17)
- 4 Roadmap: Co-evolving Brain Models, Harnesses, Data, and Ecosystems, including 4.1–4.7 (pp. 17–25)
- 5 Conclusion (p. 25)
- References (pp. 25–31)
Outside the original text pass
- Text extraction does not reconstruct figure images; inspect the retained PDF for figures and equation/table layout.
- Separate supplemental material availability has not been fully verified.
- Identity notes: the inspected title page states arXiv:2607.11689v1 [cs.RO], 13 July 2026. Its title and five authors match the catalog after name-order normalization. Only this supplied edition was reviewed; no other revision was supplied or compared.
- The catalog affiliation field contains abstract prose. The title page has a TeleAI logo but no explicit author-to-institution mapping; verified affiliation metadata is therefore omitted.
- The manifest notes that text extraction does not reconstruct figure images. This gap was addressed by inspecting the retained PDF: all 16 figures, all nine tables and all six final crops were viewed.
- Separate supplemental material availability has not been fully verified; no separate supplement was supplied. No appendix appears in the supplied 31-page PDF.
- All 11 text chunks were read individually, including the references. Full-page visual inspection covered pp. 1, 3–5 and 7–25; pp. 2, 6 and 26–31 were read as text only.
- Cited external papers, datasets, code and websites were not independently inspected. No implementation was run and no experiment was reproduced. Resource counts and descriptions below are attributed to this survey.
The visual audit above records the subsequent illustrated pass.
8.4 Traceable evidence
e-identityPDF p. 1, title/author block and left-margin arXiv stamp
The exact observed title is From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence. Authors are Yuanzhi Liang, Xufeng Zhan, Haibin Huang, Chi Zhang and Xuelong Li. The stamp identifies arXiv:2607.11689v1 [cs.RO], 13 July 2026; a TeleAI logo appears without an author-affiliation mapping.
Go to primary source ↓e-contributionsPDF p. 4, Contributions list and Section 2 opening
The paper lists scaling diagnosis, a WAM prediction contract, an embodied-brain system model, experience/evaluation contracts and a co-evolution roadmap. It describes converging research trajectories rather than a mandatory sequence of model generations.
Go to primary source ↓e-gapsPDF p. 13, Section 3.1; p. 17, Table 5 and Section 3.5
Model/representation, objective/standardization and ecosystem/systems gaps are coupled through model variables, data/task semantics and execution assumptions. Table 5 maps each gap to a proposed contract and observable test.
Go to primary source ↓e-role-taxonomyPDF p. 7, Figure 4 panels (a)–(b), caption and Section 2.5
System responsibility is separated from explicit forecasts, predictive latents and latent transition/action interfaces. The caption permits overlap and rejects a linear historical ordering; WAMs are predictive prototypes, not complete embodied brains.
Go to primary source ↓e-control-interfacesPDF p. 9, Table 1 (UniPi, VPP, LAPA and other rows), dagger footnote, caption and Section 2.6 continuation
UniPi uses inverse dynamics from visual transitions; VPP conditions inverse dynamics on predictive features; LAPA grounds discrete inter-frame latent actions with labeled robot data. Interface choice is orthogonal to modular versus joint training. Evaluation scopes differ. The dagger qualifies pi 0.7 as a VLA with optional visual subgoals, not an overall explicit-foresight architecture.
Go to primary source ↓e-data-rolesPDF p. 10, Figure 6/caption and Section 2.7; p. 11, robot trajectories and simulation/evaluation resources
Data roles depend on method and learning stage. Action-free human video may have semantic annotations but no target-robot command stream. Video representations still require action grounding through robot data or other embodiment-specific mappings.
Go to primary source ↓e-data-scalePDF p. 12, Table 2, Something-Something V2, DROID, autonomous interaction and SIMPLER rows, plus caption
The survey reports 220,847 Something-Something V2 clips and DROID's 76k trajectories, 350 hours, 564 scenes and 86 tasks. Autonomous-interaction scale is not publicly specified and no standalone public dataset is claimed. SIMPLER is an evaluation framework rather than a generic fine-tuning dataset. Resource scale and availability do not imply compatible schemas or licenses.
Go to primary source ↓e-evaluation-boundaryPDF p. 11, Section 2.8; p. 12, Table 3, all rows and caption
Predictive fidelity, interface/action grounding, closed-loop utility and embodiment-grounded executability provide complementary evidence. No layer by itself establishes full brain reasoning, a universal action space or cross-platform comparability.
Go to primary source ↓e-stack-figurePDF p. 5, Figure 2, upper prediction branch, lower execution/experience loop, legend and caption
The upper branch connects physical context/intervention to a WAM and consequence estimate. The lower branch connects brain intent, harness, execution, observed outcome, Trace Card and evaluation/admission. The upper WAM-output arrow is green dashed despite the legend labeling green dashed arrows as learning feedback; the caption identifies this branch as consequence prediction.
Go to primary source ↓e-contractPDF p. 3, Introduction prediction/intent definitions; pp. 18–19, Sections 4.1–4.2
Prediction exposes intervention consequences with horizon, reference frame, uncertainty and validity. Brain intent communicates an intended transition or capability request using an open encoding. A jointly predicted action alone is insufficient. The future architecture and any shared 3D/4D representation remain open, and joint training is permitted.
Go to primary source ↓e-ownershipPDF p. 18, Table 6 (all rows), Figure 11 and Section 4.1
The brain integrates evidence and compares interventions; the harness grounds intent and coordinates execution; specialist tools/controllers perform inference and control; the environment yields outcomes; evaluation admits traces. This is functional ownership rather than a mandatory neural architecture.
Go to primary source ↓e-harnessPDF p. 20, Section 4.3, tool/tool-model distinction and capability-contract paragraphs
A tool is a capability endpoint, whereas a tool model supplies specialized perception, prediction, planning or control. The harness registry and adapters ground requests using declared inputs, coordinates, preconditions, controllable variables, outputs, timing and failure modes.
Go to primary source ↓e-harness-figurePDF p. 21, Figure 14, legend/caption and Section 4.3 tests paragraph
Intent grounding flows through entity/frame alignment, capability resolution, precondition/safety gating and controller synchronization. Outcomes return through monitoring/recovery to a Trace Card recorder; only admitted traces are proposed as learning feedback. Suggested tests include substitution, false vetoes, failure attribution and replay.
Go to primary source ↓e-pegPDF p. 22, Figure 16 and caption
A conceptual peg-driving example decomposes a brain request into hammer acquisition, alignment and constrained impact. Arm motion drives a gripper holding an external hammer, which acts on the peg/fixture. Contracts carry frames, limits and failure modes; status returns to a Trace Card. No measured execution is presented.
Go to primary source ↓e-cardsPDF p. 22, Section 4.4; p. 23, Table 7, all three rows
Embodiment, Task and Trace Cards specify body/control metadata, goals/constraints and linked decision/execution histories respectively. These semantic records preserve original artifacts and are not a mandatory new file format.
Go to primary source ↓e-learningPDF p. 23, Table 8 (all rows and caption) and Section 4.5 continuation
Seven complementary learning objectives pair signals, updated components and required gates. The evaluator admits experience for appropriate learning uses; predictive utility, grounding, substitution, recovery and safety retention should be reported separately. No fixed chronological training recipe or numeric gate settings are provided.
Go to primary source ↓e-feedback-limitsPDF p. 22, Section 4.5 opening; p. 23, Section 4.5 continuation
Physical feedback is partially observed, embodiment-dependent and costly; valid intent can accompany control failure, and invalid intent can appear successful in a forgiving task. Examples from other domains motivate tests without proving unconstrained physical self-improvement.
Go to primary source ↓e-evaluation-layersPDF p. 24, Table 9, all rows, Status column and caption
Policy/control is marked Existing; predictive reasoning and verification/safety are Both; brain–harness grounding, tool orchestration, harness/traceability and self-improvement are Proposed. Measurements include grounding, task performance, latency, false vetoes, missed hazards, replay, regression and safety retention; no result values are reported.
Go to primary source ↓e-agendaPDF pp. 24–25, Section 4.7, three near-term priorities
The proposed milestones test decision-relevant consequence contracts, grounded representations under tool/embodiment substitution, and adaptive harnesses with attributable, replayable traces. Adapter cost and residual performance loss should be reported, and updates should pass regression and safety gates.
Go to primary source ↓e-conclusionPDF p. 25, Section 5, especially final two paragraphs
The work concludes with a roadmap. Benefits from reuse, local adaptation, compatible traces and gated post-training remain hypotheses to test. It does not claim that interfaces alone produce AGI or fix a final universal representation.
Go to primary source ↓8.5 Primary sources
From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence ↗
PDF · 16,130 extracted words
Source fingerprint
a3c0a75a8f3e9d05178a2bc786b6e9a33086a4093dfbf50a96cd44f93f94d1f4