← Search

Xin Lei

10 accepted papers

2025

A Domain Adaptation Framework for Speech Recognition Systems with Only Synthetic data

ICASSP 2025accepted

We introduce DAS (Domain Adaptation with Synthetic data), a novel domain adaptation framework for pre-trained ASR model, designed to efficiently adapt to various language-defined domains without requiring any real data. In particular, DAS first prompts large language models (LLMs) to generate domain…

Cited by 0SourceScholar
2025

Directional Source Separation for Robust Speech Recognition on Smart Glasses

ICASSP 2025accepted

Modern smart glasses leverage machine learning to offer real-time transcriptions, considerably enriching human communication experiences. However, such systems frequently encounter challenges related to environmental noises, leading to decreased speech recognition. To improve voice quality, this wor…

Cited by 15SourceScholar
2024

TODM: Train Once Deploy Many Efficient Supernet-Based RNN-T Compression For On-Device ASR Models

ICASSP 2024accepted

Automatic Speech Recognition (ASR) models need to be optimized for specific hardware before they can be deployed on devices. This can be done by tuning the model’s hyperparameters or exploring variations in its architecture. Re-training and re-validating models after making these changes can be a re…

Cited by 0SourceScholar
2023

Disentangled Training with Adversarial Examples for Robust Small-Footprint Keyword Spotting

ICASSP 2023accepted

A keyword spotting (KWS) engine continuously running on the device is exposed to various speech signals that are usually unseen beforehand. It is a challenging problem to build a small-footprint and high-performing KWS model with robustness under different acoustic environments. In this paper, we ex…

Cited by 0SourceScholar
2023

SCA: Streaming Cross-Attention Alignment For Echo Cancellation

ICASSP 2023accepted

End-to-End deep learning has shown promising results for speech enhancement tasks, such as noise suppression, dereverberation, and speech separation. However, most state-of-the-art methods for echo cancellation are either classical DSP-based or hybrid DSP-ML algorithms. Components such as the delay…

Cited by 0SourceScholar
2019

Adversarial Examples for Improving End-to-end Attention-based Small-footprint Keyword Spotting

ICASSP 2019accepted

In this paper, we explore the use of adversarial examples for improving a neural network based keyword spotting (KWS) system. Specially, in our system, an effective and small-footprint attention-based neural network model is used. Adversarial example is defined as a misclassified example by a model,…

Cited by 0SourceScholar
2019

Direct Object Recognition Without Line-Of-Sight Using Optical Coherence

CVPR 2019poster

Visual object recognition under situations in which the direct line-of-sight is blocked, such as when it is occluded around the corner, is of practical importance in a wide range of applications. With coherent illumination, the light scattered from diffusive walls forms speckle patterns that contain…

Cited by 45PDFScholar
2019

Knowledge Distillation for Recurrent Neural Network Language Modeling with Trust Regularization

ICASSP 2019accepted

Recurrent Neural Networks (RNNs) have dominated language modeling because of their superior performance over traditional N-gram based models. In many applications, a large Recurrent Neural Network language model (RNNLM) or an ensemble of several RNNLMs is used. These models have large memory footpri…

Cited by 0SourceScholar