← Search

Olivier Siohan

9 accepted papers

2024

Conformer is All You Need for Visual Speech Recognition

ICASSP 2024accepted

Visual speech recognition models extract visual features in a hierarchical manner. At the lower level, there is a visual front-end with a limited temporal receptive field that processes the raw pixels depicting the lips or faces. At the higher level, there is an encoder that attends to the embedding…

Cited by 0SourceScholar
2024

Large Scale Self-Supervised Pretraining for Active Speaker Detection

ICASSP 2024accepted

In this work we investigate the impact of a large-scale self-supervised pretraining strategy for active speaker detection (ASD) on an unlabeled dataset consisting of over 125k hours of YouTube videos. When compared to a baseline trained from scratch on much smaller in-domain labeled datasets we show…

Cited by 0SourceScholar
2022

Best of Both Worlds: Multi-Task Audio-Visual Automatic Speech Recognition and Active Speaker Detection

ICASSP 2022accepted

Under noisy conditions, automatic speech recognition (ASR) can greatly benefit from the addition of visual signals coming from a video of the speaker’s face. However, when multiple candidate speakers are visible this traditionally requires solving a separate problem, namely active speaker detection…

Cited by 0SourceScholar
2021

A Closer Look at Audio-Visual Multi-Person Speech Recognition and Active Speaker Selection

ICASSP 2021accepted

Audio-visual automatic speech recognition is a promising approach to robust ASR under noisy conditions. However, up until recently it had been traditionally studied in isolation assuming the video of a single speaking face matches the audio, and selecting the active speaker at inference time when mu…

Cited by 0SourceScholar
2020

End-to-End Multi-Person Audio/Visual Automatic Speech Recognition

ICASSP 2020accepted

Traditionally, audio-visual automatic speech recognition has been studied under the assumption that the speaking face on the visual signal is the face matching the audio. However, in a more realistic setting, when multiple faces are potentially on screen one needs to decide which face to feed to the…

Cited by 0SourceScholar
2016

Selection and combination of hypotheses for dialectal speech recognition

ICASSP 2016accepted

While research has often shown that building dialect-specific Automatic Speech Recognizers is the optimal approach to dealing with dialectal variations of the same language, we have observed that dialect-specific recognizers do not always output the best recognitions. Often enough, another dialectal…

Cited by 0SourceScholar
2015

Exemplar-based large vocabulary speech recognition using k-nearest neighbors

ICASSP 2015accepted

This paper describes a large scale exemplar-based acoustic modeling approach for large vocabulary continuous speech recognition. We construct an index of labeled training frames using high-level features extracted from the bottleneck layer of a deep neural network as indexing features. At recognitio…

Cited by 0SourceScholar