← Search

Lantian Li

10 accepted papers

2026

MT-HUBERT: SELF-SUPERVISED MIX-TRAINING FOR FEW-SHOT KEYWORD SPOTTING IN MIXED SPEECH

ICASSP 2026oral

Few-shot keyword spotting aims to detect previously unseen keywords with very limited labeled samples. A pre-training and adaptation paradigm is typically adopted for this task. While effective in clean conditions, most existing approaches struggle with mixed keyword spotting--detecting multiple ove…

Cited by 0SourcePDFScholar
2024

An Investigation of Distribution Alignment in Multi-Genre Speaker Recognition

ICASSP 2024accepted

Multi-genre speaker recognition is becoming increasingly popular due to its ability to better represent the complexities of real-world applications. However, a major challenge is the significant shift in the distribution of speaker vectors across different genres. While distribution alignment is a c…

Cited by 0SourceScholar
2021

Squeezing Value of Cross-Domain Labels: A Decoupled Scoring Approach for Speaker Verification

ICASSP 2021accepted

Domain mismatch often occurs in real applications and causes serious performance reduction on speaker verification systems. The common wisdom is to collect cross-domain data and train a multi-domain PLDA model, with the hope to learn a domain-independent speaker subspace. In this paper, we firstly p…

Cited by 0SourceScholar
2020

CN-Celeb: A Challenging Chinese Speaker Recognition Dataset

ICASSP 2020accepted

Recently, researchers set an ambitious goal of conducting speaker recognition in unconstrained conditions where the variations on ambient, channel and emotion could be arbitrary. However, most publicly available datasets are collected under constrained environments, i.e., with little noise and limit…

Cited by 0SourceScholar
2018

Human and Machine Speaker Recognition Based on Short Trivial Events

ICASSP 2018accepted

Human speech often has events that we will call trivial events, e.g., cough, laugh and sniff. Compared to regular speech, these trivial events are usually short and variable, thus generally regarded as not speaker discriminative and so are largely ignored by present speaker recognition research. How…

Cited by 0SourceScholar
2017

Speaker segmentation using deep speaker vectors for fast speaker change scenarios

ICASSP 2017accepted

A novel speaker segmentation approach based on deep neural network is proposed and investigated. This approach uses deep speaker vectors (d-vectors) to represent speaker characteristics and to find speaker change points. The d-vector is a kind of frame-level speaker discriminative feature, whose dis…

Cited by 0SourceScholar