← Search

Yake Wei

11 accepted papers

2026

Information-Theoretic Decomposition for Multimodal Interaction Learning

CVPR 2026

Multimodal learning hinges on capturing redundant, unique, and synergistic information across modalities, which collectively constitute multimodal interactions. A critical yet underexplored challenge is that these implicit interactions vary dynamically across samples. In this work, we present the fi

Cited by 0SourcecodeScholar
2025

Adaptive Unimodal Regulation for Balanced Multimodal Information Acquisition

CVPR 2025poster

Sensory training during the early ages is vital for human development. Inspired by this cognitive phenomenon, we observe that the early training stage is also important for the multimodal learning process, where dataset information is rapidly acquired. We refer to this stage as the prime learning wi…

2025

Patch Matters: Training-free Fine-grained Image Caption Enhancement via Local Perception

CVPR 2025poster

High-quality image captions play a crucial role in improving the performance of cross-modal applications such as text-to-image generation, text-to-video generation, and text-image retrieval. To generate long-form, high-quality captions, many recent studies have employed multimodal large language mod…

Cited by 1SourcePDFScholar
2025

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer

ICML 2025poster

Multimodal learning faces challenges in effectively fusing information from diverse modalities, especially when modality quality varies across samples. Dynamic fusion strategies, such as attention mechanism in Transformers, aim to address such challenge by adaptively emphasizing modalities based on…

2024

Enhancing Multimodal Cooperation via Sample-level Modality Valuation

CVPR 2024poster

One primary topic of multimodal learning is to jointly incorporate heterogeneous information from different modalities. However most models often suffer from unsatisfactory multimodal cooperation which cannot jointly utilize all modalities well. Some methods are proposed to identify and enhance the…

2024

Quantifying and Enhancing Multi-modal Robustness with Modality Preference

ICLR 2024poster

Multi-modal models have shown a promising capability to effectively integrate information from various sources, yet meanwhile, they are found vulnerable to pervasive perturbations, such as uni-modal attacks and missing conditions. To counter these perturbations, robust multi-modal representations ar…

2022

Balanced Multimodal Learning via On-the-Fly Gradient Modulation

CVPR 2022oral

Audio-visual learning helps to comprehensively understand the world, by integrating different senses. Accordingly, multiple input modalities are expected to boost model performance, but we actually find that they are not fully exploited even when the multi-modal model outperforms its uni-modal count…

Cited by 247PDFcodeScholar
2022

Learning To Answer Questions in Dynamic Audio-Visual Scenarios

CVPR 2022oral

In this paper, we focus on the Audio-Visual Question Answering (AVQA) task, which aims to answer questions regarding different visual objects, sounds, and their associations in videos. The problem requires comprehensive multimodal understanding and spatio-temporal reasoning over audio-visual scenes.…

Cited by 157PDFcodeScholar