ICASSP 2025accepted0 citations

Delving into Transformer-based Network Architecture for Guided Depth Super-Resolution

Xinchen Ye, Aokai Zhang, Rui Xu, Haojie Li

Abstract

Guided Depth Super-Resolution (GDSR) enhances low-resolution (LR) depth maps by leveraging high-resolution (HR) color images. The primary challenges involve achieving effective cross-modal data alignment and fusion, as well as incorporating multi-scale information within the Transformer architecture. To address these challenges, we propose a novel network architecture named DRMPNet which integrates two key components: Offset-based Detail Refinement (ODR) and Structure-guided Multi-scale Perception (SMP). ODR leverages offset calibration and window cross-attention to align and fuse LR depth maps with color images, effectively recovering local depth details. Meanwhile, SMP employs a structure generator and multi-scale cross-attention to capture scene details and structures at multiple scales, thereby enhancing the network’s contextual understanding. Extensive experiments on various benchmark datasets demonstrate the effectiveness of our method.

BibTeX
@inproceedings{icassp2025_delvingintotrans,
  title = {Delving into Transformer-based Network Architecture for Guided Depth Super-Resolution},
  author = {Xinchen Ye and Aokai Zhang and Rui Xu and Haojie Li},
  booktitle = {ICASSP 2025},
  year = {2025}
}