← Search

Guanyu Cai

5 accepted papers

2023

All in One: Exploring Unified Video-Language Pre-Training

CVPR 2023poster

Mainstream Video-Language Pre-training models consist of three parts, a video encoder, a text encoder, and a video-text fusion Transformer. They pursue better performance via utilizing heavier unimodal encoders or multimodal fusion Transformers, resulting in increased parameters with lower efficienc…

2023

Video-Text Pre-training with Learned Regions for Retrieval

AAAI 2023technical

Video-Text pre-training aims at learning transferable representations from large-scale video-text pairs via aligning the semantics between visual and textual information. State-of-the-art approaches extract visual features from raw pixels in an end-to-end fashion. However, these methods operate at f…

Cited by 9SourcePDFScholar
2022

Object-Aware Video-Language Pre-Training for Retrieval

CVPR 2022poster

Recently, by introducing large-scale dataset and strong transformer network, video-language pre-training has shown great success especially for retrieval. Yet, existing video-language transformer models do not explicitly fine-grained semantic align. In this work, we present Object-aware Transformers…

Cited by 91PDFcodeScholar
2021

Ask&Confirm: Active Detail Enriching for Cross-Modal Retrieval With Partial Query

ICCV 2021poster

Text-based image retrieval has seen considerable progress in recent years. However, the performance of existing methods suffers in real life since the user is likely to provide an incomplete description of an image, which often leads to results filled with false positives that fit the incomplete des…

Cited by 18PDFcodeScholar
2021

Dig into Multi-modal Cues for Video Retrieval with Hierarchical Alignment

IJCAI 2021poster

Multi-modal cues presented in videos are usually beneficial for the challenging video-text retrieval task on internet-scale datasets. Recent video retrieval methods take advantage of multi-modal cues by aggregating them to holistic high-level semantics for matching with text representations in a glo…

Cited by 24SourcePDFScholar