← Search

Qing-Yuan Jiang

11 accepted papers

2025

Rethinking Multimodal Learning from the Perspective of Mitigating Classification Ability Disproportion

NeurIPS 2025oral

Multimodal learning (MML) is significantly constrained by modality imbalance, leading to suboptimal performance in practice. While existing approaches primarily focus on balancing the learning of different modalities to address this issue, they fundamentally overlook the inherent disproportion in mo…

Cited by 0SourcecodeScholar
2025

Towards Equilibrium: An Instantaneous Probe-and-Rebalance Multimodal Learning Approach

IJCAI 2025

The multimodal imbalance problem has been extensively studied to prevent the undesirable scenario where multimodal performance falls below that of unimodal models. However, existing methods typically assess the strength of modalities and perform learning simultaneously under the imbalanced status. T

2024

Facilitating Multimodal Classification via Dynamically Learning Modality Gap

NeurIPS 2024poster

Multimodal learning falls into the trap of the optimization dilemma due to the modality imbalance phenomenon, leading to unsatisfactory performance in real applications. A core reason for modality imbalance is that the models of each modality converge at different rates. Many attempts naturally focu…

2024

TAI++: Text as Image for Multi-Label Image Classification by Co-Learning Transferable Prompt

IJCAI 2024poster

The recent introduction of prompt tuning based on pre-trained vision-language models has dramatically improved the performance of multi-label image classification. However, some existing strategies that have been explored still have drawbacks, i.e., either exploiting massive labeled visual data at a…

2022

SEMICON: A Learning-to-Hash Solution for Large-Scale Fine-Grained Image Retrieval

ECCV 2022poster

"In this paper, we propose Suppression-Enhancing Mask based attention and Interactive Channel transformatiON (SEMICON) to learn binary hash codes for dealing with large-scale fine-grained image retrieval tasks. In SEMICON, we first develop a suppression-enhancing mask (SEM) based attention to dynami…

2020

ExchNet: A Unified Hashing Network for Large-Scale Fine-Grained Image Retrieval

ECCV 2020poster

Retrieving content relevant images from a large-scale fine-grained dataset could suffer from intolerably slow query speed and highly redundant storage cost, due to high-dimensional real-valued embeddings which aim to distinguish subtle visual differences of fine-grained objects. In this paper, we st…

Cited by 52SourcePDFScholar
2019

SVD: A Large-Scale Short Video Dataset for Near-Duplicate Video Retrieval

ICCV 2019poster

With the explosive growth of video data in real applications, near-duplicate video retrieval (NDVR) has become indispensable and challenging, especially for short videos. However, all existing NDVR datasets are introduced for long videos. Furthermore, most of them are small-scale and lack of diversi…

Cited by 61PDFcodeScholar
2017

Deep Cross-Modal Hashing

CVPR 2017spotlight

Due to its low storage cost and fast query speed, cross-modal hashing (CMH) has been widely used for similarity search in multimedia retrieval applications. However, most existing CMH methods are based on hand-crafted features which might not be optimally compatible with the hash-code learning proce…

Cited by 929PDFcodeScholar