Can Real-Time Lipreading Improve Speech Recognition? A Systematic Exploration Using Human-Robot Interaction Data
Sander Goetzee, Yue Li, Koen V. Hindriks
Abstract
Speech recognition in Human-Robot Interaction (HRI) fully relies on audio-based Automatic Speech Recognition. However, speech recognition that relies solely on audio faces significant challenges in noisy environments and may lead to poor performance in such environments. One approach to address this is to also use lipreading in combination with traditional speech recognition. Recent work has shown that audiovisual speech recognition (AVSR) can achieve a Word Error Rate (WER) of only 0.9% on the dataset LRS3. In this paper, we assess the potential of combining audio with lipreading on a social robot platform, Pepper, which has not yet been widely tested for AVSR. Given that prior research has focused on non-robotic domains, it remains unclear whether such models can generalize well to social robot environments. We systematically evaluate and compare the performance of established offline and real-time audiovisual models with their audio-only counterparts. The experiments were conducted in both a controlled laboratory setting and a dynamic and noisy public environment. We evaluated the data using WER and also measured the inference latency of real-time models via Real-Time Factor and Words Per Second rates. The results demonstrate real-time performance for audio-only speech recognition across all latency metrics and near real-time performance for models that combine audio with lipreading. We also explored factors that might influence the inference performance of these models to understand how much video contributes to the audio. This includes factors related to (1) environmental and temporal variations, (2) model behavior, and (3) implementation choices. Our findings indicate that for now the audio-only models outperform the audiovisual models on a social robot platform, in contrast to what has been reported in the benchmarked literature. We conclude that more work is still needed to benefit from lipreading in HRI.
BibTeX
@inproceedings{iros2025_canrealtimelipre,
title = {Can Real-Time Lipreading Improve Speech Recognition? A Systematic Exploration Using Human-Robot Interaction Data},
author = {Sander Goetzee and Yue Li and Koen V. Hindriks},
booktitle = {IROS 2025},
year = {2025}
}