Personalizing Keyword Spotting with Speaker Information
Beltrán Labrador, Pai Zhu, Guanlong Zhao, Angelo Scorza Scarpati, Quan Wang, Alicia Lozano-Diez, Ignacio López-Moreno
Abstract
Keyword spotting systems often struggle to generalize to a diverse population with various accents and age groups. To address this challenge, we propose a novel approach that integrates speaker information into keyword spotting using Feature-wise Linear Modulation (FiLM), a recent method that allows models to learn from different data inputs and features. We explore both Text-Dependent and Text-Independent speaker recognition systems to extract speaker information, and we experiment on extracting this information from both the input audio and pre-enrolled user audio. Evaluating our systems on a diverse dataset, our primary approach yields a notable 2.6% relative improvement on Equal Error Rate overall, particularly improving performance by 5.9% for children under 12 years old and up to 24% for underrepresented speaker groups. Moreover, our proposed approach only requires a small 1% increase in the number of parameters, with a minimum impact on latency and computational cost, which makes it a practical solution for real-world applications.
BibTeX
@inproceedings{icassp2025_personalizingkey,
title = {Personalizing Keyword Spotting with Speaker Information},
author = {Beltrán Labrador and Pai Zhu and Guanlong Zhao and Angelo Scorza Scarpati and Quan Wang and Alicia Lozano-Diez and Ignacio López-Moreno},
booktitle = {ICASSP 2025},
year = {2025}
}