← Search

Ignacio López-Moreno

10 accepted papers

2025

Personalizing Keyword Spotting with Speaker Information

ICASSP 2025accepted

Keyword spotting systems often struggle to generalize to a diverse population with various accents and age groups. To address this challenge, we propose a novel approach that integrates speaker information into keyword spotting using Feature-wise Linear Modulation (FiLM), a recent method that allows…

Cited by 0SourceScholar
2024

FedAQT: Accurate Quantized Training with Federated Learning

ICASSP 2024accepted

Federated learning (FL) has been widely used to train neural networks with the decentralized training procedure where data is only accessed on clients’ devices for privacy preservation. However, the limited computation resources on clients’ devices prevent FL of large models. To overcome the constra…

Cited by 0SourceScholar
2023

Augmenting Transformer-Transducer Based Speaker Change Detection with Token-Level Training Loss

ICASSP 2023accepted

In this work we propose a novel token-based training strategy that improves Transformer-Transducer (T-T) based speaker change detection (SCD) performance. The conventional T-T based SCD model loss optimizes all output tokens equally. Due to the sparsity of the speaker changes in the training data, t…

Cited by 0SourceScholar
2023

Exploring Sequence-to-Sequence Transformer-Transducer Models for Keyword Spotting

ICASSP 2023accepted

In this paper, we present a novel approach to adapt a sequence-to-sequence Transformer-Transducer ASR system to the keyword spotting (KWS) task. We achieve this by replacing the keyword in the text transcription with a special token <kw> and training the system to detect the <kw> token in an audio s…

Cited by 0SourceScholar
2023

Locale Encoding for Scalable Multilingual Keyword Spotting Models

ICASSP 2023accepted

A Multilingual Keyword Spotting (KWS) system detects spoken keywords over multiple locales. Conventional monolingual KWS approaches do not scale well to multilingual scenarios because of high development/maintenance costs and lack of resource sharing. To overcome this limit, we propose two locale-co…

Cited by 0SourceScholar
2022

Turn-to-Diarize: Online Speaker Diarization Constrained by Transformer Transducer Speaker Turn Detection

ICASSP 2022accepted

In this paper, we present a novel speaker diarization system for streaming on-device applications. In this system, we use a transformer transducer to detect the speaker turns, represent each speaker turn by a speaker embedding, then cluster these embeddings with constraints from the detected speaker…

Cited by 61SourceScholar
2018

Attention-Based Models for Text-Dependent Speaker Verification

ICASSP 2018accepted

Attention-based models have recently shown great performance on a range of tasks, such as speech recognition, machine translation, and image captioning due to their ability to summarize relevant information that expands through the entire length of an input sequence. In this paper, we analyze the us…

Cited by 0SourceScholar