← Search

Nanxin Chen

10 accepted papers

2023

A Quantum Kernel Learning Approach to Acoustic Modeling for Spoken Command Recognition

ICASSP 2023accepted

We propose a quantum kernel learning (QKL) framework to address the inherent data sparsity issues often encountered in training large-scare acoustic models in low-resource scenarios. We project acoustic features based on classical-to-quantum feature encoding. Different from existing quantum convolut…

Cited by 0SourceScholar
2023

From English to More Languages: Parameter-Efficient Model Reprogramming for Cross-Lingual Speech Recognition

ICASSP 2023accepted

In this work, we propose a new parameter-efficient learning framework based on neural model reprogramming for cross-lingual speech recognition, which can re-purpose well-trained English automatic speech recognition (ASR) models to recognize the other languages. We design different auxiliary neural a…

Cited by 0SourceScholar
2021

Focus on the Present: A Regularization Method for the ASR Source-Target Attention Layer

ICASSP 2021accepted

This paper introduces a novel method to diagnose the source-target attention in state-of-the-art end-to-end speech recognition models with joint connectionist temporal classification (CTC) and attention training. Our method is based on the fact that both, CTC and source-target attention, are acting…

Cited by 0SourceScholar
2021

WaveGrad: Estimating Gradients for Waveform Generation

ICLR 2021poster

This paper introduces WaveGrad, a conditional model for waveform generation which estimates gradients of the data density. The model is built on prior work on score matching and diffusion probabilistic models. It starts from a Gaussian white noise signal and iteratively refines the signal via a grad…

2020

Feature Enhancement with Deep Feature Losses for Speaker Verification

ICASSP 2020accepted

Speaker Verification still suffers from the challenge of generalization to novel adverse environments. We leverage on the recent advancements made by deep learning based speech enhancement and propose a feature-domain supervised denoising based solution. We propose to use Deep Feature Loss which opt…

Cited by 0SourceScholar
2020

Improving Language Identification for Multilingual Speakers

ICASSP 2020accepted

Spoken language identification (LID) technologies have improved in recent years from discriminating largely distinct languages to discriminating highly similar languages or even dialects of the same language. One aspect that has been mostly neglected, however, is discrimination of languages for mult…

Cited by 0SourceScholar
2020

X-Vectors Meet Emotions: A Study On Dependencies Between Emotion and Speaker Recognition

ICASSP 2020accepted

In this work, we explore the dependencies between speaker recognition and emotion recognition. We first show that knowledge learned for speaker recognition can be reused for emotion recognition through transfer learning. Then, we show the effect of emotion on speaker recognition. For emotion recogni…

Cited by 0SourceScholar
2020

Zero-Shot Multi-Speaker Text-To-Speech with State-Of-The-Art Neural Speaker Embeddings

ICASSP 2020accepted

While speaker adaptation for end-to-end speech synthesis using speaker embeddings can produce good speaker similarity for speakers seen during training, there remains a gap for zero-shot adaptation to unseen speakers. We investigate multi-speaker modeling for end-to-end text-to-speech synthesis and…

Cited by 0SourceScholar
2018

Measuring Uncertainty in Deep Regression Models: The Case of Age Estimation from Speech

ICASSP 2018accepted

Age estimation from speech recently received a lot of attention. Approaches such as i-vectors and deep learning have been successfully applied to this task achieving great performance. However, one drawback of those methods is that they produce a hard age estimation without any kind of confidence me…

Cited by 0SourceScholar