ICML 2026poster0 citations

OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention

Zhangquan Chen, Jiale Tao, Ruihuang Li, Yihao Hu, Ruitao Chen, Zhantao Yang, Xinlei Yu, Haodong Jing

Abstract

Humans perceive the world through diverse modalities that operate synergistically to support a holistic understanding of their surroundings. However, existing omnimodal models still exhibit substantial performance degradation on visual tasks when the audio modality is incorporated. We identify this “modality interference” as a consequence of pre-training data imbalances, where the scarcity of mixed modality supervision induces a bias towards isolated modalities, resulting in an inherent trade-off. To address this challenge, we propose OmniVideo-R1, a novel reinforced reasoning framework that leverages post-training to rectify modality bias. OmniVideo-R1 empowers models to “think with omnimodal cues” and integrate cross-modal information. The framework consists of two key strategies: (1) query-intensive grounding based on self-supervised learning paradigms; and (2) modality- attentive fusion built upon contrastive learning paradigms. Extensive experiments on multiple benchmarks demonstrate that OmniVideo-R1 consistently outperforms strong baselines, highlighting its effectiveness and robust generalization capabilities.

TransformerTheoryRobustnessFairnessVisionRetrievalBenchmark
BibTeX
@inproceedings{
chen2026omnivideor,
title={OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention},
author={Zhangquan Chen and Jiale Tao and Ruihuang Li and Yihao Hu and Ruitao Chen and Zhantao Yang and Xinlei Yu and Haodong Jing and Manyuan Zhang and Shuai Shao and Biao Wang and Qinglin Lu and Ruqi Huang},
booktitle={Forty-third International Conference on Machine Learning},
year={2026},
url={https://openreview.net/forum?id=he06cvibXv}
}