← Search

Zhenyu Huang

8 accepted papers

2026

MR-RAG: Multimodal Relevance-Aware Retrieval-Augmented Generation for Medical Visual Question Answering

CVPR 2026

Large Vision Language Models (LVLMs) with retrieval-augmented generation (RAG) are emerging as a main paradigm for processing vision-language medical tasks due to their promising achievements. However, existing approaches exhibit two significant limitations in both retrieval and generation stage: Fi

Cited by 0SourceScholar
2024

Multi-granularity Correspondence Learning from Long-term Noisy Videos

ICLR 2024oral

Existing video-language studies mainly focus on learning short video clips, leaving long-term temporal dependencies rarely explored due to over-high computational cost of modeling long videos. To address this issue, one feasible solution is learning the correspondence between video clips and caption…

2023

Robust Domain Adaptation for Machine Reading Comprehension

AAAI 2023technical

Most domain adaptation methods for machine reading comprehension (MRC) use a pre-trained question-answer (QA) construction model to generate pseudo QA pairs for MRC transfer. Such a process will inevitably introduce mismatched pairs (i.e., Noisy Correspondence) due to i) the unavailable QA pairs in…

Cited by 1SourcePDFScholar
2022

Learning With Twin Noisy Labels for Visible-Infrared Person Re-Identification

CVPR 2022poster

In this paper, we study an untouched problem in visible-infrared person re-identification (VI-ReID), namely, Twin Noise Labels (TNL) which refers to as noisy annotation and correspondence. In brief, on the one hand, it is inevitable to annotate some persons with the wrong identity due to the complex…

Cited by 213PDFcodeScholar
2021

Learning with Noisy Correspondence for Cross-modal Matching

NeurIPS 2021oral

Cross-modal matching, which aims to establish the correspondence between two different modalities, is fundamental to a variety of tasks such as cross-modal retrieval and vision-and-language understanding. Although a huge number of cross-modal matching methods have been proposed and achieved remarkab…

2021

Partially View-Aligned Representation Learning With Noise-Robust Contrastive Loss

CVPR 2021poster

In real-world applications, it is common that only a portion of data is aligned across views due to spatial, temporal, or spatiotemporal asynchronism, thus leading to the so-called Partially View-aligned Problem (PVP). To solve such a less-touched problem without the help of labels, we propose simul…

Cited by 170PDFScholar
2019

COMIC: Multi-view Clustering Without Parameter Selection

ICML 2019oral

In this paper, we study two challenges in clustering analysis, namely, how to cluster multi-view data and how to perform clustering without parameter selection on cluster size. To this end, we propose a novel objective function to project raw data into one space in which the projection embraces the…

Cited by 374SourcePDFScholar