CrossHOI: Learning Cross-View Representations for Monocular 3D Human-Object Interaction Reconstruction
Reconstructing 3D human-object interaction (HOI) from monocular images is highly challenging especially when human and object are mutually occluded. Existing methods primarily rely on single-view inputs, which fundamentally limit their ability to recover occluded regions and accurately estimate cont