IROS 20250 citations

3D-AMTA: Occlusion-Aware Real-Time 3D Hand Pose Estimation with Auto Mask and Token-Specific Attention

Dongfang Zhao, Menghe Zhang, Yangwen Liang, Shuangquan Wang, Kee-Bong Song, Donghoon Kim

Abstract

Understanding hand motion from a single RGB image is challenging due to occlusions and high articulation. This paper presents 3D-AMTA, a transformer-based framework with Auto Mask and Token-specific Attention for occlusion-aware 3D hand pose estimation (HPE). We propose two novel architectural enhancements: auto mask for high-occlusion scenarios, and token-specific attention for fine-grained hand articulations. These modules seamlessly integrate into transformer-based architectures that enhance real-time performance in interactive systems. To enable efficient deployment on robotic and embedded platforms, we propose 3D-AMTA-Mobile, a lightweight variant optimized for on-device processing. It achieves 267 FPS on NVIDIA RTX 2080Ti-GPU while maintaining high accuracy, making it well-suited for resource-constrained robotic applications. Extensive evaluations on FreiHAND and HO3D demonstrate that our approach consistently outperforms state-of-the-art methods in terms of accuracy, efficiency, and inference speed. These advancements contribute to robust hand perception for interactive robotics and AR-based teleoperation.

BibTeX
@inproceedings{iros2025_3damtaocclusiona,
  title = {3D-AMTA: Occlusion-Aware Real-Time 3D Hand Pose Estimation with Auto Mask and Token-Specific Attention},
  author = {Dongfang Zhao and Menghe Zhang and Yangwen Liang and Shuangquan Wang and Kee-Bong Song and Donghoon Kim},
  booktitle = {IROS 2025},
  year = {2025}
}
3D-AMTA: Occlusion-Aware Real-Time 3D Hand Pose Estimation with Auto Mask and Token-Specific Attention · IROS 2025