← Search

Naomi Harte

5 accepted papers

2025

Visual Cues Enhance Predictive Turn-Taking for Two-Party Human Interaction

ACL 2025finding

Turn-taking is richly multimodal. Predictive turn-taking models (PTTMs) facilitate natural- istic human-robot interaction, yet most rely solely on speech. We introduce MM-VAP, a multimodal PTTM which combines speech with visual cues including facial expression, head pose and gaze. We find that it ou…

2023

Learnable Frontends That Do Not Learn: Quantifying Sensitivity To Filterbank Initialisation

ICASSP 2023accepted

While much of modern speech and audio processing relies on deep neural networks trained using fixed audio representations, recent studies suggest great potential in acoustic frontends learnt jointly with a backend. In this study, we focus specifically on learnable filterbanks. Prior studies have rep…

Cited by 0SourceScholar
2020

Cogans For Unsupervised Visual Speech Adaptation To New Speakers

ICASSP 2020accepted

Audio-Visual Speech Recognition (AVSR) faces the difficult task of exploiting acoustic and visual cues simultaneously. Augmenting speech with the visual channel creates its own challenges, e.g. every person has unique mouth movements, making the generalization of visual models very difficult. This f…

Cited by 0SourceScholar