Exo2EgoHOI: Hand-Object-Interaction Aware Exocentric-to-Egocentric Video Generation

Hongjia Zhai1, Xiyu Zhang2, Haoran Zhang1, Zhichao Ye3, Haomin Liu3, Guofeng Zhang2,3, Ian Reid1, Xingxing Zuo1

1MBZUAI2Zhejiang University3InSpatio

Overview

Exocentric human demonstration is translated by Exo2EgoHOI into egocentric views and retargeted to a robot executing the same manipulation.
Exo2EgoHOI translates exocentric human demonstrations into egocentric observations that can guide robot execution through retargeting.

Abstract

Egocentric videos of human manipulation provide valuable visual experience for embodied intelligence, yet collecting such data at scale is costly. Exocentric-to-egocentric video generation offers a scalable alternative by transforming abundant third-person manipulation videos into first-person observations. However, existing methods often struggle to faithfully preserve demonstrated hand-object interactions (HOI) across large viewpoint changes due to insufficient fine-grained interaction guidance and weak object-centric anchoring. We present Exo2EgoHOI, an HOI-aware video generative framework for interaction-preserving exocentric-to-egocentric translation. To preserve fine-grained HOI, we introduce a unified 4D HOI prior that combines scene geometry, articulated hand renderings, and dense hand-object relation fields, together with a dual-branch residual adapter for injecting structural and relational cues into the video generation backbone. To preserve object consistency, we introduce Decomposed Gated Cross-Attention, which separately encodes object and background references and adaptively integrates global semantic and local appearance features as object-centric anchors. Experiments on ARCTIC-HOI and Ego-Exo4D demonstrate substantial improvements in object consistency and HOI preservation while maintaining competitive visual fidelity. In particular, on ARCTIC-HOI, Exo2EgoHOI improves object mIoU by 32.3% and reduces MPJPE and PA-MPJPE by 34.7% and 50.0%, respectively, relative to the respective best baseline results.

Method overview

Exo2EgoHOI pipeline: exocentric video, world-space 4D HOI, unified scene, hand, and interaction priors, HOI adapters, object and background anchors, video DiT, and egocentric output.
A unified 4D HOI prior guides scene geometry, hand motion, and interactions. HOI adapters inject these cues into the video model, while Decomposed Gated Cross-Attention uses object and background references to preserve appearance.

Quantitative comparisons

Full test sets. Best values are bold; second-best values are underlined.

ARCTIC-HOI

ARCTIC-HOI quantitative results. Higher is better for PSNR, SSIM, CLIP-I, mIoU, and CLIP-O; lower is better for the remaining metrics.
MethodVisual fidelityObject consistencyHand consistency
PSNR ↑SSIM ↑LPIPS ↓CLIP-I ↑mIoU ↑Center-Err. ↓CLIP-O ↑MPJPE ↓WA-MPJPE ↓PA-MPJPE ↓
WAN VACE12.780.61630.51150.84850.23360.08930.80802870.22371.1820.36
EgoWorld11.900.46520.62030.79950.22410.11310.78692282.03440.6057.92
Vista4D13.580.66140.43240.87370.27680.09290.81762123.78387.7328.91
EgoX11.830.59410.47430.89100.29330.08040.83452019.81325.5622.12
Ours13.430.66850.38600.91750.38810.05830.85571318.74313.7510.18

Ego-Exo4D

Ego-Exo4D quantitative results. Higher is better for PSNR, SSIM, CLIP-I, mIoU, and CLIP-O; lower is better for the remaining metrics.
MethodVisual fidelityObject consistencyHand consistency
PSNR ↑SSIM ↑LPIPS ↓CLIP-I ↑mIoU ↑Center-Err. ↓CLIP-O ↑MPJPE ↓WA-MPJPE ↓PA-MPJPE ↓
WAN VACE12.670.41610.59810.82510.03620.22140.80036810.622163.1874.00
EgoWorld11.050.34100.71860.72200.02270.23490.7911———
Vista4D10.460.26840.73480.74590.01540.25710.806720074.432293.86120.68
EgoX13.140.44740.52580.85630.07030.18420.82985773.681928.0775.55
Ours14.110.46900.50350.86650.07230.17450.83415208.171749.3693.81