← Search

Yan-Bo Lin

6 accepted papers

2023

Vision Transformers Are Parameter-Efficient Audio-Visual Learners

CVPR 2023poster

Vision transformers (ViTs) have achieved impressive results on various computer vision tasks in the last several years. In this work, we study the capability of frozen ViTs, pretrained only on visual data, to generalize to audio-visual data without finetuning any of its original parameters. To do so…

2022

ECLIPSE: Efficient Long-Range Video Retrieval Using Sight and Sound

ECCV 2022poster

"We introduce an audiovisual method for long-range text-to-video retrieval. Unlike previous approaches designed for short video retrieval (e.g., 5-15 seconds in duration), our approach aims to retrieve minute-long videos that capture complex human actions. One challenge of standard video-only approa…

2021

Exploiting Audio-Visual Consistency with Partial Supervision for Spatial Audio Generation

AAAI 2021technical

Human perceives rich auditory experience with distinct sound heard by ears. Videos recorded with binaural audio particular simulate how human receives ambient sound. However, a large number of videos are with monaural audio only, which would degrade the user experience due to the lack of ambient inf…

Cited by 24SourcePDFScholar
2021

Exploring Cross-Video and Cross-Modality Signals for Weakly-Supervised Audio-Visual Video Parsing

NeurIPS 2021poster

The audio-visual video parsing task aims to temporally parse a video into audio or visual event categories. However, it is labor intensive to temporally annotate audio and visual events and thus hampers the learning of a parsing model. To this end, we propose to explore additional cross-video and cr…

2019

Cross-Dataset Person Re-Identification via Unsupervised Pose Disentanglement and Adaptation

ICCV 2019poster

Person re-identification (re-ID) aims at recognizing the same person from images taken across different cameras. To address this challenging task, existing re-ID models typically rely on a large amount of labeled training data, which is not practical for real-world applications. To alleviate this li…

Cited by 248PDFScholar