← Search

Sourish Chaudhuri

3 accepted papers

2020

Ava Active Speaker: An Audio-Visual Dataset for Active Speaker Detection

ICASSP 2020accepted

Active speaker detection is an important component in video analysis algorithms for applications such as speaker diarization, video re-targeting for meetings, speech enhancement, and human-robot interaction. The absence of a large, carefully labeled audio-visual active speaker dataset has limited ev…

Cited by 0SourceScholar
2018

Using audio-visual information to understand speaker activity: Tracking active speakers on and off screen

ICASSP 2018accepted

We present a system that associates faces with voices in a video by fusing information from the audio and visual signals. The thesis underlying our work is that an extreme simple approach to generating (weak) speech clusters can be combined with strong visual signals to effectively associate faces a…

Cited by 10SourceScholar
2017

CNN architectures for large-scale audio classification

ICASSP 2017accepted

Convolutional Neural Networks (CNNs) have proven very effective in image classification and show promise for audio. We use various CNN architectures to classify the soundtracks of a dataset of 70M training videos (5.24 million hours) with 30,871 video-level labels. We examine fully connected Deep Ne…

Cited by 3037SourceScholar