← Search

Bangxiang Lan

4 accepted papers

2026

SAVE: Speech-Aware Video Representation Learning for Video-Text Retrieval

CVPR 2026

For video-text retrieval, the use of CLIP has been a de facto standard. However, as CLIP provides only image and text encoders, this consensus has led to a biased paradigm that entirely ignores the sound track of videos. While several attempts have been made to reintroduce audio -- typically by inco

Cited by 0SourcecodeScholar
2025

Hybrid-Tower: Fine-grained Pseudo-query Interaction and Generation for Text-to-Video Retrieval

ICCV 2025accepted

The Text-to-Video Retrieval (T2VR) task aims to retrieve unlabeled videos by textual queries with the same semantic meanings. Recent CLIP-based approaches have explored two frameworks: Two-Tower versus Single-Tower framework, yet the former suffers from low effectiveness, while the latter suffers fr…

Cited by 0SourcePDFScholar
2025

Multi-Object Sketch Animation by Scene Decomposition and Motion Planning

ICCV 2025poster

Sketch animation, which brings static sketches to life by generating dynamic video sequences, has found widespread applications in GIF design, cartoon production, and daily entertainment. While current methods for sketch animation perform well in single-object sketch animation, they struggle in mult…

Cited by 0SourcePDFScholar
2024

Holistic Features are almost Sufficient for Text-to-Video Retrieval

CVPR 2024poster

For text-to-video retrieval (T2VR) which aims to retrieve unlabeled videos by ad-hoc textual queries CLIP-based methods currently lead the way. Compared to CLIP4Clip which is efficient and compact state-of-the-art models tend to compute video-text similarity through fine-grained cross-modal feature…