CVPR 20260 citations

Anchor-Guided Gradient Alignment for Incomplete Multimodal Learning

Zhi-Hao Guan, Longfei Huang, Yang Yang

Abstract

Vision-language pre-training (VLP) has achieved remarkable performance across diverse multimodal learning (MML) tasks. Recently, many efforts have focused on reconstructing missing modalities to improve the adaptability of VLP models in incomplete MML scenarios. However, these approaches overlook the learning imbalance under severe missing-modality conditions, i.e., the optimization process is dominated by reconstructed samples, thereby weakening complete-sample representations. In this paper, we propose a novel ANchor-guided Gradient Alignment (ANGA) framework to address this issue. Specifically, we first retrieve similar instances to reconstruct the missing modalities, thereby alleviating information deficiency. We then introduce an entropy-driven curriculum that progressively incorporates reliable reconstructed samples together with complete ones to form an optimization anchor, which guides gradient alignment to mitigate learning imbalance. Furthermore, we design a semantic-enhanced adapter that leverages the retrieved instances to generate dynamic prompts, further enhancing the robustness of the VLP model. Extensive experiments on widely used datasets demonstrate the superiority of ANGA over state-of-the-art (SOTA) baselines across various missing-modality scenarios. The code is available at this repository.

BibTeX
@inproceedings{cvpr2026_anchorguidedgrad,
  title = {Anchor-Guided Gradient Alignment for Incomplete Multimodal Learning},
  author = {Zhi-Hao Guan and Longfei Huang and Yang Yang},
  booktitle = {CVPR 2026},
  year = {2026}
}
Anchor-Guided Gradient Alignment for Incomplete Multimodal Learning · CVPR 2026