Being-H0.7: A Latent World-Action Model from Egocentric Videos
Classification
View four quadrants- Major category
- WAMs
- Quadrant
- Not assigned
- Architecture
- Not assigned
- Prediction paradigm
- Not assigned
- Subcategories
- Latent prediction & JEPA
- Source review status
- Not assigned
Category review. A shared prior/posterior policy trains current-context latent queries to match future-informed representations, and those anticipatory queries guide action generation. This is an implicit latent world-action mechanism, with no explicit video rollout or IDM at inference. Reading evidence
Contribution
Being-H0.7 trains a robot policy to anticipate useful future information inside latent queries. A future-aware training branch aligns its hidden states with a deployable branch that sees only current context, avoiding visual rollout during control. Strong benchmark and real-robot results support the complete system, while missing component ablations leave the causal contribution of future alignment unresolved (E02–E09, E12, E15).
Abstract
An abstract has not been added yet.
Affiliations
Not listed in the collection.