← Search

Varsha Hedau

3 accepted papers

2022

Asd-Transformer: Efficient Active Speaker Detection Using Self And Multimodal Transformers

ICASSP 2022accepted

Multimodal active speaker detection (ASD) methods assign a speaking/not-speaking label per individual in a video clip. ASD is critical for applications such as natural human-computer interaction, speaker diarization, and video reframing. Recent work has shown the success of transformers in multimoda…

Cited by 0SourceScholar
2022

FashionVLP: Vision Language Transformer for Fashion Retrieval With Feedback

CVPR 2022poster

Fashion image retrieval based on a query pair of reference image and natural language feedback is a challenging task that requires models to assess fashion related information from visual and textual modalities simultaneously. We propose a new vision-language transformer based model, FashionVLP, tha…

Cited by 120PDFScholar