← Search

Kong Aik Lee

29 accepted papers

2026

ADDRESSING GRADIENT MISALIGNMENT IN DATA-AUGMENTED TRAINING FOR ROBUST SPEECH DEEPFAKE DETECTION

ICASSP 2026oral

In speech deepfake detection (SDD), data augmentation (DA) is commonly used to improve model generalization across varied speech conditions and spoofing attacks. However, during training, the backpropagated gradients from original and augmented inputs may misalign, which can result in conflicting pa…

Cited by 0SourcePDFScholar
2026

SPEAKING CLEARLY: A SIMPLIFIED WHISPER-BASED CODEC FOR LOW-BITRATE SPEECH CODING

ICASSP 2026poster

Speech codecs serve as bridges between continuous speech signals and large language models, yet face an inherent conflict between acoustic fidelity and semantic preservation. To mitigate this conflict, prevailing methods augment acoustic codecs with complex semantic supervision. We explore the oppos…

Cited by 0SourcePDFScholar
2026

STREAM-VOICE-ANON: ENHANCING UTILITY OF REAL-TIME SPEAKER ANONYMIZATION VIA NEURAL AUDIO CODEC AND LANGUAGE MODELS

ICASSP 2026poster

Protecting speaker identity is crucial for online voice applications, yet streaming speaker anonymization (SA) remains underexplored. Recent research has demonstrated that neural audio codec (NAC) provides superior speaker feature disentanglement and linguistic fidelity. NAC can also be used with ca…

Cited by 0SourcePDFScholar
2025

Grouped Knowledge Distillation with Adaptive Logit Softening for Speaker Recognition

ICASSP 2025accepted

Recent works suggest that decoupling the information of non-target speakers from that of the target speaker in knowledge distillation (KD) and subsequently emphasizing the former can lead to significant performance improvement. However, a well-trained teacher model typically produces almost zero non…

Cited by 0SourceScholar
2025

LlamaPartialSpoof: An LLM-Driven Fake Speech Dataset Simulating Disinformation Generation

ICASSP 2025accepted

Previous fake speech datasets were constructed from a defender’s perspective to develop countermeasure (CM) systems without considering diverse motivations of attackers. To better align with real-life scenarios, we created LlamaPartialSpoof, a 130-hour dataset that contains both fully and partially…

Cited by 0SourceScholar
2025

Text-dependent Speaker Verification Challenge 2024: Exploring Shared and User-defined Passphrases

ICASSP 2025accepted

In contrast to text-independent speaker verification, which has received significant attention from researchers and has many competitions dedicated to it, text-dependent speaker verification (TdSV) has been less explored recently. The TdSV Challenge 2024 was organized to analyze and explore novel me…

Cited by 0SourceScholar
2024

CPAUG: Refining Copy-Paste Augmentation for Speech Anti-Spoofing

ICASSP 2024accepted

Conventional copy-paste augmentations generate new training instances by concatenating existing utterances to increase the amount of data for neural network training. However, the direct application of copy-paste augmentation for anti-spoofing is problematic. This paper refines the copy-paste augmen…

Cited by 0SourceScholar
2024

Emphasized Non-Target Speaker Knowledge in Knowledge Distillation for Automatic Speaker Verification

ICASSP 2024accepted

Knowledge distillation (KD) is used to enhance automatic speaker verification performance by ensuring consistency between large teacher networks and lightweight student networks at the embedding level or label level. However, the conventional label-level KD overlooks the significant knowledge from n…

Cited by 0SourceScholar
2024

Gradient Weighting for Speaker Verification in Extremely Low Signal-to-Noise Ratio

ICASSP 2024accepted

Speaker verification is hampered by background noise, particularly at extremely low Signal-to-Noise Ratio (SNR) under 0 dB. It is difficult to suppress noise without introducing unwanted artifacts, which adversely affects speaker verification. We proposed the mechanism called Gradient Weighting (Gra…

Cited by 0SourceScholar
2024

Two-stage Semi-supervised Speaker Recognition with Gated Label Learning

IJCAI 2024poster

Speaker recognition technologies have been successfully applied in diverse domains, benefiting from the advance of deep learning. Nevertheless, current efforts are still subject to the lack of labeled data. Such issues have been attempted in computer vision, through semi-supervised learning (SSL) th…

Cited by 1SourcePDFScholar
2023

Cross-Modal Audio-Visual Co-Learning for Text-Independent Speaker Verification

ICASSP 2023accepted

Visual speech (i.e., lip motion) is highly related to auditory speech due to the co-occurrence and synchronization in speech production. This paper investigates this correlation and proposes a cross-modal speech co-learning paradigm. The primary motivation of our cross-modal co-learning method is mo…

Cited by 0SourceScholar
2023

Disentangling Voice and Content with Self-Supervision for Speaker Recognition

NeurIPS 2023poster

For speaker recognition, it is difficult to extract an accurate speaker representation from speech because of its mixture of speaker traits and content. This paper proposes a disentanglement framework that simultaneously models speaker traits and content variability in speech. It is realized with t…

Cited by 51SourcePDFScholar
2023

Incorporating Uncertainty from Speaker Embedding Estimation to Speaker Verification

ICASSP 2023accepted

Speech utterances recorded under differing conditions exhibit varying degrees of confidence in their embedding estimates, i.e., uncertainty, even if they are extracted using the same neural network. This paper aims to incorporate the uncertainty estimate produced in the xi-vector network front-end w…

Cited by 0SourceScholar
2023

Leveraging Positional-Related Local-Global Dependency for Synthetic Speech Detection

ICASSP 2023accepted

Automatic speaker verification (ASV) systems are vulnerable to spoofing attacks. As synthetic speech exhibits local and global artifacts compared to natural speech, incorporating local-global dependency would lead to better anti-spoofing performance. To this end, we propose the Rawformer that levera…

Cited by 0SourceScholar
2023

Noise-Disentanglement Metric Learning for Robust Speaker Verification

ICASSP 2023accepted

Automatic speaker verification (ASV) suffers from performance degradation in noisy environments. To solve this problem, we propose the noise-disentanglement metric learning to reduce the speaker-irrelevant noisy components and build a noise-invariant embedding space. Specifically, the disentanglemen…

Cited by 0SourceScholar
2023

Probabilistic Back-ends for Online Speaker Recognition and Clustering

ICASSP 2023accepted

This paper focuses on multi-enrollment speaker recognition which naturally occurs in the task of online speaker clustering, and studies the properties of different scoring back-ends in this scenario. First, we show that popular cosine scoring suffers from poor score calibration with a varying number…

Cited by 0SourceScholar
2023

Self-Supervised Audio-Visual Speaker Representation with Co-Meta Learning

ICASSP 2023accepted

In self-supervised speaker verification, the quality of pseudo labels determines the upper bound of its performance and it is not uncommon to end up with massive amount of unreliable pseudo labels. We observe that the complementary information in different modalities ensures a robust supervisory sig…

Cited by 0SourceScholar
2022

Improving Contextual Coherence in Variational Personalized and Empathetic Dialogue Agents

ICASSP 2022accepted

In recent years, latent variable models, such as the Conditional Variational Auto Encoder (CVAE), have been applied to both personalized and empathetic dialogue generation. Prior work have largely focused on generating diverse dialogue responses that exhibit persona consistency and empathy. However,…

Cited by 0SourceScholar
2022

Learning Domain-Invariant Transformation for Speaker Verification

ICASSP 2022accepted

Automatic speaker verification (ASV) faces domain shift caused by the mismatch of intrinsic and extrinsic factors such as recording device and speaking style in real-world applications, which leads to unsatisfactory performance. To this end, we propose the meta generalized transformation via meta-le…

Cited by 0SourceScholar
2022

MFA: TDNN with Multi-Scale Frequency-Channel Attention for Text-Independent Speaker Verification with Short Utterances

ICASSP 2022accepted

The time delay neural network (TDNN) represents one of the state-of-the-art of neural solutions to text-independent speaker verification. However, they require a large number of filters to capture the speaker characteristics at any local frequency region. In addition, the performance of such systems…

Cited by 0SourceScholar
2022

Self-Supervised Speaker Recognition with Loss-Gated Learning

ICASSP 2022accepted

In self-supervised learning for speaker recognition, pseudo labels are useful as the supervision signals. It is a known fact that a speaker recognition model doesn’t always benefit from pseudo labels due to their unreliability. In this work, we observe that a speaker recognition network tends to mod…

Cited by 0SourceScholar
2022

Summary on the ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Grand Challenge

ICASSP 2022accepted

The ICASSP 2022 Multi-channel Multi-party Meeting Transcription Grand Challenge (M2MeT) focuses on one of the most valuable and the most challenging scenarios of speech technologies. The M2MeT challenge has particularly set up two tracks, speaker diarization (track 1) and multi-speaker automatic spe…

Cited by 0SourceScholar
2021

COOPNet: Multi-Modal Cooperative Gender Prediction in Social Media User Profiling

ICASSP 2021accepted

The principal way of performing user profiling is to investigate accumulated social media data. However, the problem of information asymmetry generally exists in user generated contents since users post multi-modal contents in social media freely. In this paper, we propose a novel text-image coopera…

Cited by 0SourceScholar
2021

Meta-Learning for Cross-Channel Speaker Verification

ICASSP 2021accepted

Automatic speaker verification (ASV) has been successfully deployed for identity recognition. With increasing use of ASV technology in real-world applications, channel mismatch caused by the recording devices and environments severely degrade its performance, especially in the case of unseen channel…

Cited by 0SourceScholar
2021

Replay-Attack Detection Using Features With Adaptive Spectro-Temporal Resolution

ICASSP 2021accepted

Variable-resolution processing aims to improve the feature representation ability by enlarging the local discriminative details. In previous anti-spoofing studies, different phones and frequency regions were both proven to have various levels of sensitivity to replay distortion. In this paper, an ad…

Cited by 0SourceScholar
2020

A Generalized Framework for Domain Adaptation of PLDA in Speaker Recognition

ICASSP 2020accepted

This paper proposes a generalized framework for domain adaptation of Probabilistic Linear Discriminant Analysis (PLDA) in speaker recognition. It not only includes several existing supervised and unsupervised domain adaptation methods but also makes possible more flexible usage of available data in…

Cited by 0SourceScholar