Trust, but Verify: Uncertainty-Driven Evidential Multimodal Representation Learning
Yupeng Han, Kai Zhang, Xianquan Wang, Zhihong Pan, Ze Liu, Jiyuan He
Abstract
Effective multimodal learning in real-world scenarios depends on a nuanced treatment of uncertainty, which arises at three levels: (1) Intrinsic Uncertainty from modality-specific noise or ambiguity; (2) Relational Uncertainty due to cross-modal conflicts or redundancy; and (3) Aggregated Uncertainty when fusing potentially inconsistent signals. Most existing methods overlook this hierarchy, applying a single uncertainty model. We propose Adaptive Evidential Multimodal Representation Learning (AEMRL), a framework aligned with this multi-level view. To address intrinsic uncertainty, Disentangled Evidential Uncertainty Encoding (D-EUE) provides interpretable, class-aware reliability scores per modality. For relational uncertainty, Uncertainty-Conditioned Dynamic Factorization (UDF) uses a hypernetwork to dynamically extract complementary cues and suppress conflict. To resolve aggregated uncertainty, Adaptive Fusion & Conflict-Aware Calibration (AFCAC) adaptively weights evidence streams and calibrates final predictions based on detected conflicts. Extensive experiments on five diverse benchmark datasets show that AEMRL consistently enhances task accuracy, reduces calibration error, and improves robustness to noise, semantic conflict, and missing modalities.
BibTeX
@inproceedings{ijcai2026_trustbutverifyun,
title = {Trust, but Verify: Uncertainty-Driven Evidential Multimodal Representation Learning},
author = {Yupeng Han and Kai Zhang and Xianquan Wang and Zhihong Pan and Ze Liu and Jiyuan He},
booktitle = {IJCAI 2026},
year = {2026}
}