← Search

Guanlong Zhao

7 accepted papers

2025

Personalizing Keyword Spotting with Speaker Information

ICASSP 2025accepted

Keyword spotting systems often struggle to generalize to a diverse population with various accents and age groups. To address this challenge, we propose a novel approach that integrates speaker information into keyword spotting using Feature-wise Linear Modulation (FiLM), a recent method that allows…

Cited by 0SourceScholar
2024

USM-SCD: Multilingual Speaker Change Detection Based on Large Pretrained Foundation Models

ICASSP 2024accepted

We introduce a multilingual speaker change detection model (USM-SCD) that can simultaneously detect speaker turns and perform ASR for 96 languages. This model is adapted from a speech foundation model trained on a large quantity of supervised and unsupervised data, demonstrating the utility of fine-…

Cited by 0SourceScholar
2023

Augmenting Transformer-Transducer Based Speaker Change Detection with Token-Level Training Loss

ICASSP 2023accepted

In this work we propose a novel token-based training strategy that improves Transformer-Transducer (T-T) based speaker change detection (SCD) performance. The conventional T-T based SCD model loss optimizes all output tokens equally. Due to the sparsity of the speaker changes in the training data, t…

Cited by 0SourceScholar
2023

Exploring Sequence-to-Sequence Transformer-Transducer Models for Keyword Spotting

ICASSP 2023accepted

In this paper, we present a novel approach to adapt a sequence-to-sequence Transformer-Transducer ASR system to the keyword spotting (KWS) task. We achieve this by replacing the keyword in the text transcription with a special token <kw> and training the system to detect the <kw> token in an audio s…

Cited by 0SourceScholar
2018

Accent Conversion Using Phonetic Posteriorgrams

ICASSP 2018accepted

Accent conversion (AC) aims to transform non-native speech to sound as if the speaker had a native accent. This can be achieved by mapping source spectra from a native speaker into the acoustic space of the non-native speaker. In prior work, we proposed an AC approach that matches frames between the…

Cited by 0SourceScholar
2018

Voice Conversion Through Residual Warping in a Sparse, Anchor-Based Representation of Speech

ICASSP 2018accepted

In previous work we presented a Sparse, Anchor-Based Representation of speech (SABR) that uses phonemic “anchors” to represent an utterance with a set of sparse non-negative weights. SABR is speaker-independent: combining weights from a source speaker with anchors from a target speaker can be used f…

Cited by 0SourceScholar