← Search

Yuqing Song

3 accepted papers

2023

Accommodating Audio Modality in CLIP for Multimodal Processing

AAAI 2023technical

Multimodal processing has attracted much attention lately especially with the success of pre-training. However, the exploration has mainly focused on vision-language pre-training, as introducing more modalities can greatly complicate model design and optimization. In this paper, we extend the state-…

2022

Unifying Event Detection and Captioning as Sequence Generation via Pre-training

ECCV 2022poster

"Dense video captioning aims to generate corresponding text descriptions for a series of events in the untrimmed video, which can be divided into two sub-tasks, event detection and event captioning. Unlike previous works that tackle the two sub-tasks separately, recent works have focused on enhancin…