← Search

Otavio Braga

4 accepted papers

2024

Large Scale Self-Supervised Pretraining for Active Speaker Detection

ICASSP 2024accepted

In this work we investigate the impact of a large-scale self-supervised pretraining strategy for active speaker detection (ASD) on an unlabeled dataset consisting of over 125k hours of YouTube videos. When compared to a baseline trained from scratch on much smaller in-domain labeled datasets we show…

Cited by 0SourceScholar
2022

Best of Both Worlds: Multi-Task Audio-Visual Automatic Speech Recognition and Active Speaker Detection

ICASSP 2022accepted

Under noisy conditions, automatic speech recognition (ASR) can greatly benefit from the addition of visual signals coming from a video of the speaker’s face. However, when multiple candidate speakers are visible this traditionally requires solving a separate problem, namely active speaker detection…

Cited by 0SourceScholar
2021

A Closer Look at Audio-Visual Multi-Person Speech Recognition and Active Speaker Selection

ICASSP 2021accepted

Audio-visual automatic speech recognition is a promising approach to robust ASR under noisy conditions. However, up until recently it had been traditionally studied in isolation assuming the video of a single speaking face matches the audio, and selecting the active speaker at inference time when mu…

Cited by 0SourceScholar
2020

End-to-End Multi-Person Audio/Visual Automatic Speech Recognition

ICASSP 2020accepted

Traditionally, audio-visual automatic speech recognition has been studied under the assumption that the speaking face on the visual signal is the face matching the audio. However, in a more realistic setting, when multiple faces are potentially on screen one needs to decide which face to feed to the…

Cited by 0SourceScholar