← Search

Boyun Zhang

1 accepted papers

2026

HVD: HUMAN VISION-DRIVEN VIDEO REPRESENTATION LEARNING FOR TEXT-VIDEO RETRIEVAL

ICASSP 2026poster

The success of CLIP has driven substantial progress in text-video retrieval. However, current methods often suffer from "blind" feature interaction, where the model struggles to discern key visual information from background noise due to the sparsity of textual queries. To bridge this gap, we draw i…

Cited by 0SourcePDFScholar