← Search

Luhui Xu

2 accepted papers

2023

Token Mixing: Parameter-Efficient Transfer Learning from Image-Language to Video-Language

AAAI 2023technical

Applying large scale pre-trained image-language model to video-language tasks has recently become a trend, which brings two challenges. One is how to effectively transfer knowledge from static images to dynamic videos, and the other is how to deal with the prohibitive cost of fully fine-tuning due t…

2022

TS2-Net: Token Shift and Selection Transformer for Text-Video Retrieval

ECCV 2022poster

"Text-Video retrieval is a task of great practical value and has received increasing attention, among which learning spatial-temporal video representation is one of the research hotspots. The video encoders in the state-of-the-art video retrieval models usually directly adopt the pre-trained vision…