Open-Modality Latent Modality Interaction Maximization for Audio-Visual Learning
Zhao Yang, Rui Jiang, Xiao Fu, Wei Xi, Jizhong Zhao
Abstract
The utilization of multimodal cues enhances the effectiveness of specific cognitive tasks in audio-visual learning. However, on the one hand, designing a unified model for multimodal learning poses challenges due to the presence of information redundancy and modality noise. On the other hand, existing multimodal models face limitations in handling the modality-missing inference. In this work, we propose a Latent Modality Interaction with mutual information Maximization (LMIM) model architecture for multimodal learning, which effectively integrates multimodal cues by learning essential modality information and reducing the redundant information. We employ a group of latent tokens as pivots to filter out noise and redundancy across different modalities. Simultaneously, mutual information maximization and distribution alignment are utilized to preserve task-related information through multimodal fusion. Furthermore, a random modality masking training strategy is employed to mitigate potential over-reliance on dominant modality. Extensive experiments demonstrate that our model achieves significant improvement over current competitive baselines on two datasets, including UCF51 and Kinetics-Sounds datasets.
BibTeX
@inproceedings{icassp2025_openmodalitylate,
title = {Open-Modality Latent Modality Interaction Maximization for Audio-Visual Learning},
author = {Zhao Yang and Rui Jiang and Xiao Fu and Wei Xi and Jizhong Zhao},
booktitle = {ICASSP 2025},
year = {2025}
}