AAAI 2026technical0 citations

SAME: Spatial-Aware Multimodal Egocentric Human Pose Estimation

Yurong Fu, Peng Dai, Yu Zhang, Feng Yiqiang, Yang Zhang, Haoqian Wang

Abstract

Egocentric human pose estimation (HPE) plays a crucial role in immersive applications such as virtual and augmented reality. However, existing methods relying on either visual or sparse inertial data alone often suffer from occlusion or ill-posed problems. In this work, we propose SAME, a novel spatial-aware multimodal fusion framework combining the complementary signals from the stereo images and sparse IMUs for accurate and robust egocentric HPE. It adopts a two-stage network based on a dual coordinate frame to mitigate the coordinate inconsistencies among the stereo cameras and the IMUs. In the first stage, the IMU signals are transformed into the local frame and iteratively fused with the stereo images for estimating 3D poses in the local frame. In the second stage, the local poses are transformed into the global frame with the 6DOF head poses provided by the head-mounted display

BibTeX
@inproceedings{aaai2026_samespatialaware,
  title = {SAME: Spatial-Aware Multimodal Egocentric Human Pose Estimation},
  author = {Yurong Fu and Peng Dai and Yu Zhang and Feng Yiqiang and Yang Zhang and Haoqian Wang},
  booktitle = {AAAI 2026},
  year = {2026}
}