CVPR 20260 citations

AIMDepth: Asymmetric Image-Event Mamba for Monocular Depth Estimation

Luoxi Jing, Dianxi Shi, Yushe Cao, Yuanze Wang, Junze Zhang, Yuning Cui, Mengzhu Wang

Abstract

Monocular depth estimation is essential for applications such as robotics. The complementary characteristics of event and image modalities have inspired fusion-based methods for robust depth estimation. However, existing methods rely on convolutional or attention-based architectures, which either have limited capacity for long-range modeling or incur high computational cost, making them less suitable for depth estimation over long sequences. Moreover, effective image-event fusion remains challenging, since most methods directly fuse features without addressing the domain gap and representational differences between raw events and images, resulting in semantic bias and degraded performance. In this work, we propose AIMDepth, an Asymmetric Image-Event Mamba framework for monocular depth estimation, built on state space models for linear complexity and accurate prediction. To alleviate input-domain misalignment, we introduce a Spectral Cross-modal Prior Guidance (SCPG) for bidirectional prior injection at the input level. To reduce the imbalance between sparse events and dense images, we design an asymmetric modal-aware Encoder (AME) with separate encoding paths and feature-level alignment. We further develop a Modality-interactive Local Refinement (ModiLocal) to enable hierarchical interaction and fine-grained alignment. Experiments on public datasets show that AIMDepth achieves state-of-the-art performance in complex environments.

BibTeX
@inproceedings{cvpr2026_aimdepthasymmetr,
  title = {AIMDepth: Asymmetric Image-Event Mamba for Monocular Depth Estimation},
  author = {Luoxi Jing and Dianxi Shi and Yushe Cao and Yuanze Wang and Junze Zhang and Yuning Cui and Mengzhu Wang},
  booktitle = {CVPR 2026},
  year = {2026}
}
AIMDepth: Asymmetric Image-Event Mamba for Monocular Depth Estimation · CVPR 2026