← Search

Yanxiong Li

8 accepted papers

2025

Cross-Domain Few-Shot Open-Set Keyword Spotting Using Keyword Adaptation and Prototype Reprojection

ICASSP 2025accepted

Personalized keyword spotting (KWS) with few enrollment utterances remains an important problem over years. KWS remains a challenging task due to the following factors, including the scarcity of enrollment samples, speech variation in the open-set scenarios, and distributional gap between source and…

Cited by 0SourceScholar
2025

Deep Enhancement Spotting Network for Low-complexity Keyword Spotting in Noisy Environments

ICASSP 2025accepted

Keyword Spotting (KWS) is crucial for hands-free voice-activated systems, requiring a balance between accuracy and complexity, especially in noisy environments. While Speech Enhancement (SE) can improve KWS accuracy, existing methods often lack the ability to effectively utilize the rich features pr…

Cited by 0SourceScholar
2023

Clean Sample Guided Self-Knowledge Distillation for Image Classification

ICASSP 2023accepted

For two-stage knowledge distillation, the combination with Data Augmentation (DA) is straightforward and effective. Yet, for online Self-knowledge Distillation (SD), DA is not always beneficial because of the absence of a trustworthy teacher model. To address this issue, this paper proposes an SD me…

Cited by 0SourceScholar
2021

A Stage Match for Query-by-Example Spoken Term Detection Based On Structure Information of Query

ICASSP 2021accepted

The state-of-the-art of query-by-example spoken term detection (QbE-STD) strategies are usually based on segmental dynamic time warping (S-DTW). However, the sliding window in S-DTW may separate signal of a word into different segments and produce many illegal candidates required to be compared with…

Cited by 0SourceScholar
2021

Domestic Activities Clustering From Audio Recordings Using Convolutional Capsule Autoencoder Network

ICASSP 2021accepted

Recent efforts have been made on domestic activities classification from audio recordings, especially the works submitted to the challenge of DCASE (Detection and Classification of Acoustic Scenes and Events) since 2018. In contrast, few studies were done on domestic activities clustering, which is…

Cited by 0SourceScholar
2020

Sound Event Detection Via Dilated Convolutional Recurrent Neural Networks

ICASSP 2020accepted

Convolutional recurrent neural networks (CRNNs) have achieved state-of-the-art performance for sound event detection (SED). In this paper, we propose to use a dilated CRNN, namely a CRNN with a dilated convolutional kernel, as the classifier for the task of SED. We investigate the effectiveness of d…

Cited by 54SourceScholar
2017

Mobile phone clustering from acquired speech recordings using deep Gaussian supervector and spectral clustering

ICASSP 2017accepted

Acquisition device clustering from speech recordings is a new and critical problem in the field of speech forensic, which aims at merging speech recordings acquired by the same device into one cluster without both pre-knowing prior information of the processed data and pre-training classifier. We pr…

Cited by 0SourceScholar
2016

Source cell phone matching from speech recordings by sparse representation and KISS metric

ICASSP 2016accepted

Source recording device matching from two speech recordings is a new and important problem of digital media forensics. It aims to answer the question that whether or not two speech recordings are recorded by the same recording device. In this study we propose a source cell phone matching scheme. The…

Cited by 0SourceScholar