← Search

Tan Yu

15 accepted papers

2026

MoRA: Missing Modality Low-Rank Adaptation for Visual Recognition

ICLR 2026poster

Pre-trained vision language models have shown remarkable performance on visual recognition tasks, but they typically assume the availability of complete multimodal inputs during both training and inference. In real-world scenarios, however, modalities may be missing due to privacy constraints, colle…

Cited by 0SourcecodeScholar
2025

KALAHash: Knowledge-Anchored Low-Resource Adaptation for Deep Hashing

AAAI 2025technical

Deep hashing has been widely used for large-scale approximate nearest neighbor search due to its storage and search efficiency. However, existing deep hashing methods predominantly rely on abundant training data, leaving the more challenging scenario of low-resource adaptation for deep hashing relat…

2022

Cross-Lingual Cross-Modal Consolidation for Effective Multilingual Video Corpus Moment Retrieval

NAACL 2022findings

Existing multilingual video corpus moment retrieval (mVCMR) methods are mainly based on a two-stream structure. The visual stream utilizes the visual content in the video to estimate the query-visual similarity, and the subtitle stream exploits the query-subtitle similarity. The final query-video si…

Cited by 23SourcePDFScholar
2021

Inflate and Shrink:Enriching and Reducing Interactions for Fast Text-Image Retrieval

EMNLP 2021main

By exploiting the cross-modal attention, cross-BERT methods have achieved state-of-the-art accuracy in cross-modal retrieval. Nevertheless, the heavy text-image interactions in the cross-BERT model are prohibitively slow for large-scale retrieval. Late-interaction methods trade off retrieval accurac…

Cited by 18SourcePDFScholar
2019

Temporal Structure Mining for Weakly Supervised Action Detection

ICCV 2019poster

Different from the fully-supervised action detection problem that is dependent on expensive frame-level annotations, weakly supervised action detection (WSAD) only needs video-level annotations, making it more practical for real-world applications. Existing WSAD methods detect action instances by sc…

Cited by 95PDFScholar
2017

HOPE: Hierarchical Object Prototype Encoding for Efficient Object Instance Search in Videos

CVPR 2017poster

This paper tackles the problem of efficient and effective object instance search in videos. To effectively capture the relevance between a query and video frames and precisely localize the particular object, we leverage the object proposals to improve the quality of object instance search in videos.…

Cited by 16PDFScholar