← Search

Eng Siong Chng

59 accepted papers

2026

ADDRESSING GRADIENT MISALIGNMENT IN DATA-AUGMENTED TRAINING FOR ROBUST SPEECH DEEPFAKE DETECTION

ICASSP 2026oral

In speech deepfake detection (SDD), data augmentation (DA) is commonly used to improve model generalization across varied speech conditions and spoofing attacks. However, during training, the backpropagated gradients from original and augmented inputs may misalign, which can result in conflicting pa…

Cited by 0SourcePDFScholar
2026

ALIGNING GENERATIVE SPEECH ENHANCEMENT WITH PERCEPTUAL FEEDBACK

ICASSP 2026oral

Language Model (LM)-based speech enhancement (SE) has recently emerged as a promising direction, but existing approaches predominantly rely on token-level likelihood objectives that weakly reflect human perception. This mismatch limits progress, as optimizing signal accuracy does not always improve…

Cited by 0SourcePDFScholar
2026

AudioTrust: Benchmarking The Multifaceted Trustworthiness of Audio Large Language Models

ICLR 2026poster

The rapid development and widespread adoption of Audio Large Language Models (ALLMs) require a rigorous assessment of their trustworthiness. However, existing evaluation frameworks, primarily designed for text, are not equipped to handle the unique vulnerabilities introduced by audio’s acoustic prop…

Cited by 0SourcecodeScholar
2026

Omni-Captioner: Data Pipeline, Models, and Benchmark for Omni Detailed Perception

ICLR 2026poster

Fine-grained perception of multimodal information is critical for advancing human–AI interaction. With recent progress in audio–visual technologies, Omni Language Models (OLMs), capable of processing audio and video signals in parallel, have emerged as a promising paradigm for achieving richer unde…

Cited by 0SourcecodeScholar
2026

STREAM-VOICE-ANON: ENHANCING UTILITY OF REAL-TIME SPEAKER ANONYMIZATION VIA NEURAL AUDIO CODEC AND LANGUAGE MODELS

ICASSP 2026poster

Protecting speaker identity is crucial for online voice applications, yet streaming speaker anonymization (SA) remains underexplored. Recent research has demonstrated that neural audio codec (NAC) provides superior speaker feature disentanglement and linguistic fidelity. NAC can also be used with ca…

Cited by 0SourcePDFScholar
2026

SUMMARY ON THE MULTILINGUAL CONVERSATIONAL SPEECH LANGUAGE MODEL CHALLENGE: DATASETS, TASKS, BASELINES, AND METHODS

ICASSP 2026poster

This paper summarizes the Interspeech2025 Multilingual Conversational Speech Language Model (MLC-SLM) challenge, which aims to advance the exploration of building effective multilingual conversational speech LLMs (SLLMs). We provide a detailed description of the task settings for the MLC-SLM challen…

Cited by 0SourcePDFScholar
2026

SumRA: Parameter Efficient Fine-tuning with Singular Value Decomposition and Summed Orthogonal Basis

ICLR 2026poster

Parameter-efficient fine-tuning (PEFT) aims to adapt large pretrained speech models using fewer trainable parameters while maintaining performance. Low-Rank Adaptation (LoRA) achieves this by decomposing weight updates into two low-rank matrices, $A$ and $B$, such that $W'=W_0+BA$. Previous studies…

Cited by 0SourceScholar
2025

Extending Whisper for Emotion Prediction Using Word-level Pseudo Labels

ICASSP 2025accepted

This paper extends Whisper’s automatic speech recognition (ASR) capabilities to perform speech-based emotion recognition (SER) by incorporating word-level emotion classification alongside ASR output. We generate four emotion pseudo-labels (neutral, happy, sad, angry) for each word using a pretrained…

Cited by 0SourceScholar
2025

Intra-modal and Cross-modal Synchronization for Audio-visual Deepfake Detection and Temporal Localization

ICCV 2025poster

Recent deepfake detection algorithms focus solely on uni-modal or cross-modal inconsistencies. While the former disregards audio-visual correspondence entirely rendering them less effective against multimodal attacks, the latter overlooks inconsistencies in a particular modality. Moreover, many mode…

2025

LlamaPartialSpoof: An LLM-Driven Fake Speech Dataset Simulating Disinformation Generation

ICASSP 2025accepted

Previous fake speech datasets were constructed from a defender’s perspective to develop countermeasure (CM) systems without considering diverse motivations of attackers. To better align with real-life scenarios, we created LlamaPartialSpoof, a 130-hour dataset that contains both fully and partially…

Cited by 26SourceScholar
2025

Speech Enhancement Using Continuous Embeddings of Neural Audio Codec

ICASSP 2025accepted

Recent advancements in Neural Audio Codec (NAC) models have inspired their use in various speech processing tasks, including speech enhancement (SE). In this work, we propose a novel, efficient SE approach by leveraging the pre-quantization output of a pretrained NAC encoder. Unlike prior NAC-based…

Cited by 0SourceScholar
2024

Are Soft Prompts Good Zero-Shot Learners for Speech Recognition?

ICASSP 2024accepted

Large self-supervised pre-trained speech models require computationally expensive fine-tuning for downstream tasks. Soft prompt tuning offers a simple parameter-efficient alternative by utilizing minimal soft prompt guidance, enhancing portability while also maintaining competitive performance. Howe…

Cited by 0SourceScholar
2024

Cross-Modality and Within-Modality Regularization for Audio-Visual Deepfake Detection

ICASSP 2024accepted

Audio-visual deepfake detection scrutinizes manipulations in public video using complementary multimodal cues. Current methods, which train on fused multimodal data for multimodal targets face challenges due to uncertainties and inconsistencies in learned representations caused by independent modali…

Cited by 0SourceScholar
2024

Emphasized Non-Target Speaker Knowledge in Knowledge Distillation for Automatic Speaker Verification

ICASSP 2024accepted

Knowledge distillation (KD) is used to enhance automatic speaker verification performance by ensuring consistency between large teacher networks and lightweight student networks at the embedding level or label level. However, the conventional label-level KD overlooks the significant knowledge from n…

Cited by 0SourceScholar
2024

Enhancing Low-Latency Speaker Diarization with Spatial Dictionary Learning

ICASSP 2024accepted

This study proposes a low-latency online speaker diarization framework. Specifically, we design a spatial dictionary learning module shared across different frequency bands, enabling spatial feature learning at each frequency bin. This contributes to reducing the latency constraints of the online di…

Cited by 0SourceScholar
2024

GenTranslate: Large Language Models are Generative Multilingual Speech and Machine Translators

ACL 2024long

Recent advances in large language models (LLMs) have stepped forward the development of multilingual speech and machine translation by its reduced representation errors and incorporated external knowledge. However, both translation tasks typically utilize beam search decoding and top-1 hypothesis se…

2024

Noise-Aware Speech Separation with Contrastive Learning

ICASSP 2024accepted

Recently, speech separation (SS) task has achieved remarkable progress driven by deep learning technique. However, it is still challenging to separate target speech from noisy mixture, as the neural model is vulnerable to assign background noise to each speaker. In this paper, we propose a noise-awa…

Cited by 0SourceScholar
2024

SPGM: Prioritizing Local Features for Enhanced Speech Separation Performance

ICASSP 2024accepted

Dual-path is a popular architecture for speech separation models (e.g. Sepformer) which splits long sequences into overlapping chunks for its intra- and inter-blocks that separately model intra-chunk local features and inter-chunk global relationships. However, it has been found that inter-blocks, w…

Cited by 0SourceScholar
2023

Contrastive Speech Mixup for Low-Resource Keyword Spotting

ICASSP 2023accepted

Most of the existing neural-based models for keyword spotting (KWS) in smart devices require thousands of training samples to learn a decent audio representation. However, with the rising demand for smart devices to become more person-alized, KWS models need to adapt quickly to smaller user samples.…

Cited by 0SourceScholar
2023

Cross-Modal Global Interaction and Local Alignment for Audio-Visual Speech Recognition

IJCAI 2023poster

Audio-visual speech recognition (AVSR) research has gained a great success recently by improving the noise-robustness of audio-only automatic speech recognition (ASR) with noise-invariant visual information. However, most existing AVSR approaches simply fuse the audio and visual features by concaten…

2023

De'hubert: Disentangling Noise in a Self-Supervised Model for Robust Speech Recognition

ICASSP 2023accepted

Existing self-supervised pre-trained speech models have offered an effective way to leverage massive unannotated corpora to build good automatic speech recognition (ASR). However, many current models are trained on a clean corpus from a single source, which tends to do poorly when noise is present d…

Cited by 0SourceScholar
2023

Gradient Remedy for Multi-Task Learning in End-to-End Noise-Robust Speech Recognition

ICASSP 2023accepted

Speech enhancement (SE) is proved effective in reducing noise from noisy speech signals for downstream automatic speech recognition (ASR), where multi-task learning strategy is employed to jointly optimize these two tasks. However, the enhanced speech learned by SE objective may not always yield goo…

Cited by 0SourceScholar
2023

Hearing Lips in Noise: Universal Viseme-Phoneme Mapping and Transfer for Robust Audio-Visual Speech Recognition

ACL 2023long

Audio-visual speech recognition (AVSR) provides a promising solution to ameliorate the noise-robustness of audio-only speech recognition with visual information. However, most existing efforts still focus on audio modality to improve robustness considering its dominance in AVSR task, with noise adap…

2023

Improving Spoken Language Identification with Map-Mix

ICASSP 2023accepted

The pre-trained multi-lingual XLSR model generalizes well for language identification after fine-tuning on unseen languages. However, the performance significantly degrades when the languages are not very distinct from each other, for example, in the case of dialects. Low resource dialect classifica…

Cited by 0SourceScholar
2023

Leveraging Modality-Specific Representations for Audio-Visual Speech Recognition via Reinforcement Learning

AAAI 2023technical

Audio-visual speech recognition (AVSR) has gained remarkable success for ameliorating the noise-robustness of speech recognition. Mainstream methods focus on fusing audio and visual inputs to obtain modality-invariant representations. However, such representations are prone to over-reliance on audio…

Cited by 31SourcePDFScholar
2023

MIR-GAN: Refining Frame-Level Modality-Invariant Representations with Adversarial Network for Audio-Visual Speech Recognition

ACL 2023long

Audio-visual speech recognition (AVSR) attracts a surge of research interest recently by leveraging multimodal signals to understand human speech. Mainstream approaches addressing this task have developed sophisticated architectures and techniques for multi-modality fusion and representation learnin…

2023

Metric-Oriented Speech Enhancement Using Diffusion Probabilistic Model

ICASSP 2023accepted

Deep neural network based speech enhancement technique focuses on learning a noisy-to-clean transformation supervised by paired training data. However, the task-specific evaluation metric (e.g., PESQ) is usually non-differentiable and can not be directly constructed in the training criteria. This mi…

Cited by 0SourceScholar
2023

Probabilistic Back-ends for Online Speaker Recognition and Clustering

ICASSP 2023accepted

This paper focuses on multi-enrollment speaker recognition which naturally occurs in the task of online speaker clustering, and studies the properties of different scoring back-ends in this scenario. First, we show that popular cosine scoring suffers from poor score calibration with a varying number…

Cited by 5SourceScholar
2023

Speech-Text Based Multi-Modal Training with Bidirectional Attention for Improved Speech Recognition

ICASSP 2023accepted

To let the state-of-the-art end-to-end ASR model enjoy data efficiency, as well as much more unpaired text data by multi-modal training, one needs to address two problems: 1) the synchronicity of feature sampling rates between speech and language (aka text data); 2) the homogeneity of the learned re…

Cited by 0SourceScholar
2023

UniS-MMC: Multimodal Classification via Unimodality-supervised Multimodal Contrastive Learning

ACL 2023findings

Multimodal learning aims to imitate human beings to acquire complementary information from multiple modalities for various downstream tasks. However, traditional aggregation-based multimodal fusion methods ignore the inter-modality relationship, treat each modality equally, suffer sensor noise, and…

2023

Unifying Speech Enhancement and Separation with Gradient Modulation for End-to-End Noise-Robust Speech Separation

ICASSP 2023accepted

Recent studies in neural network-based monaural speech separation (SS) have achieved a remarkable success thanks to increasing ability of long sequence modeling. However, they would degrade significantly when put under realistic noisy conditions, as the background noise could be mistaken for speaker…

Cited by 0SourceScholar
2022

An Embarrassingly Simple Model for Dialogue Relation Extraction

ICASSP 2022accepted

Dialogue relation extraction (RE) is to predict the relation type of two entities mentioned in a dialogue. In this paper, we propose a simple yet effective model named SimpleRE for the RE task. SimpleRE captures the interrelations among multiple relations in a dialogue through a novel input format n…

Cited by 0SourceScholar
2022

Convmixer: Feature Interactive Convolution with Curriculum Learning for Small Footprint and Noisy Far-Field Keyword Spotting

ICASSP 2022accepted

Building efficient architecture in neural speech processing is paramount to success in keyword spotting deployment. However, it is very challenging for lightweight models to achieve noise robustness with concise neural operations. In a real-world application, the user environment is typically noisy…

Cited by 0SourceScholar
2022

Interactive Feature Fusion for End-to-End Noise-Robust Speech Recognition

ICASSP 2022accepted

Speech enhancement (SE) aims to suppress the additive noise from noisy speech signals to improve the speech’s perceptual quality and intelligibility. However, the over-suppression phenomenon in the enhanced speech might degrade the performance of downstream automatic speech recognition (ASR) task du…

Cited by 0SourceScholar
2022

L-SpEx: Localized Target Speaker Extraction

ICASSP 2022accepted

Speaker extraction aims to extract the target speaker’s voice from a multi-talker speech mixture given an auxiliary reference utterance. Recent studies show that speaker extraction benefits from the location or direction of the target speaker. However, these studies assume that the target speaker’s…

Cited by 0SourceScholar
2022

Minimum Word Error Training For Non-Autoregressive Transformer-Based Code-Switching ASR

ICASSP 2022accepted

Non-autoregressive end-to-end ASR framework might be potentially appropriate for code-switching recognition task thanks to its inherent property that present output token being independent of historical ones. However, it still under-performs the state-of-the-art autoregressive ASR frameworks. In thi…

Cited by 0SourceScholar
2022

Noise-Robust Speech Recognition With 10 Minutes Unparalleled In-Domain Data

ICASSP 2022accepted

Noise-robust speech recognition systems require large amounts of training data including noisy speech data and corresponding transcripts to achieve state-of-the-art performances in face of various practical environments. However, such plenty of in-domain data is not always available in the real-life…

Cited by 0SourceScholar
2022

Self-Critical Sequence Training for Automatic Speech Recognition

ICASSP 2022accepted

Although automatic speech recognition (ASR) task has gained remarkable success by sequence-to-sequence models, there are two main mismatches between its training and testing that might lead to performance degradation: 1) The typically used cross-entropy criterion aims to maximize log-likelihood of t…

Cited by 0SourceScholar
2022

Speech Emotion Recognition with Co-Attention Based Multi-Level Acoustic Information

ICASSP 2022accepted

Speech Emotion Recognition (SER) aims to help the machine to understand human’s subjective emotion from only audio in-formation. However, extracting and utilizing comprehensive in-depth audio information is still a challenging task. In this paper, we propose an end-to-end speech emotion recognition…

Cited by 0SourceScholar
2021

A Unified Speaker Adaptation Approach for ASR

EMNLP 2021main

Transformer models have been used in automatic speech recognition (ASR) successfully and yields state-of-the-art results. However, its performance is still affected by speaker mismatch between training and test data. Further finetuning a trained model with target speaker data is the most natural app…

2021

GDPNet: Refining Latent Multi-View Graph for Relation Extraction

AAAI 2021technical

Relation Extraction (RE) is to predict the relation type of two entities that are mentioned in a piece of text, e.g., a sentence or a dialogue. When the given text is long, it is challenging to identify indicative words for the relation prediction. Recent advances on RE task are from BERT-based sequ…

2021

Learning Disentangled Feature Representations for Speech Enhancement Via Adversarial Training

ICASSP 2021accepted

Neural speech enhancement degrades significantly in face of unseen noise. To address such mismatch, we propose to learn noise-agnostic feature representations by disentanglement learning, which removes the unspecified noise factor, while keeping the specified factors of variation associated with the…

Cited by 0SourceScholar
2021

Multi-Stage Speaker Extraction with Utterance and Frame-Level Reference Signals

ICASSP 2021accepted

Speaker extraction requires a sample speech from the target speaker as the reference. However, enrolling a speaker with a long speech is not practical. We propose a speaker extraction technique, that performs in multiple stages to take full advantage of short reference speech sample. The extracted s…

Cited by 0SourceScholar
2021

Preventing Early Endpointing for Online Automatic Speech Recognition

ICASSP 2021accepted

With the recent development of end-to-end models in speech recognition, there have been more interests in adapting these models for online speech recognition. However, using end-to-end models for online speech recognition is known to suffer from an early endpointing problem, which brings in many del…

Cited by 0SourceScholar
2021

Representation Learning with Spectro-Temporal-Channel Attention for Speech Emotion Recognition

ICASSP 2021accepted

Convolutional neural network (CNN) is found to be effective in learning representation for speech emotion recognition. CNNs do not explicitly model the associations or relative importance of features in the spectral/temporal/channel-wise axes. In this paper, we propose an attention module, named spe…

Cited by 0SourceScholar
2020

Independent Language Modeling Architecture for End-To-End ASR

ICASSP 2020accepted

The attention-based end-to-end (E2E) automatic speech recognition (ASR) architecture allows for joint optimization of acoustic and language models within a single network. However, in a vanilla E2E ASR architecture, the decoder sub-network (subnet), which incorporates the role of the language model…

Cited by 0SourceScholar
2020

Time-Domain Neural Network Approach for Speech Bandwidth Extension

ICASSP 2020accepted

In this paper, we study the time-domain neural network approach for speech bandwidth extension. We propose a network architecture, named multi-scale fusion neural network (MfNet), that gradually restores the low-frequency signal and predicts the high-frequency signal through the exchange of informat…

Cited by 0SourceScholar
2019

Optimization of Speaker Extraction Neural Network with Magnitude and Temporal Spectrum Approximation Loss

ICASSP 2019accepted

The SpeakerBeam-FE (SBF) method is proposed for speaker extraction. It attempts to overcome the problem of unknown number of speakers in an audio recording during source separation. The mask approximation loss of SBF is sub-optimal, which doesn't calculate direct signal reconstruction error and cons…

Cited by 0SourceScholar
2018

Single Channel Speech Separation with Constrained Utterance Level Permutation Invariant Training Using Grid LSTM

ICASSP 2018accepted

Utterance level permutation invariant training (uPIT) technique is a state-of-the-art deep learning architecture for speaker independent multi-talker separation. uPIT solves the label ambiguity problem by minimizing the mean square error (MSE) over all permutations between outputs and targets. Howev…

Cited by 0SourceScholar
2018

Unsupervised Domain Adaptation via Domain Adversarial Training for Speaker Recognition

ICASSP 2018accepted

The i-vector approach to speaker recognition has achieved good performance when the domain of the evaluation dataset is similar to that of the training dataset. However, in realworld applications, there is always a mismatch between the training and evaluation datasets, that leads to performance degr…

Cited by 0SourceScholar
2017

On time-frequency mask estimation for MVDR beamforming with application in robust speech recognition

ICASSP 2017accepted

Acoustic beamforming has played a key role in the robust automatic speech recognition (ASR) applications. Accurate estimates of the speech and noise spatial covariance matrices (SCM) are crucial for successfully applying the minimum variance distortionless response (MVDR) beamforming. Reliable estim…

Cited by 0SourceScholar
2016

An expectation-maximization eigenvector clustering approach to direction of arrival estimation of multiple speech sources

ICASSP 2016accepted

This paper presents an eigenvector clustering approach for estimating the direction of arrival (DOA) of multiple speech signals using a microphone array. Existing clustering approaches usually only use low frequencies to avoid spatial aliasing. In this study, we propose a probabilistic eigenvector c…

Cited by 0SourceScholar
2016

Approximate search of audio queries by using DTW with phone time boundary and data augmentation

ICASSP 2016accepted

Dynamic Time Warping (DTW) is widely used in language independent query-by-example (QbE) spoken term detection (STD) tasks due to its high performance. However, there are two limitations of DTW based template matching, 1) it is not straightforward to perform approximate match of audio queries; 2) DT…

Cited by 0SourceScholar
2016

Combining non-negative matrix factorization and deep neural networks for speech enhancement and automatic speech recognition

ICASSP 2016accepted

Sparse Non-negative Matrix Factorization (SNMF) and Deep Neural Networks (DNN) have emerged individually as two efficient machine learning techniques for single-channel speech enhancement. Nevertheless, there are only few works investigating the combination of SNMF and DNN for speech enhancement and…

Cited by 0SourceScholar
2016

Content-aware local variability vector for speaker verification with short utterance

ICASSP 2016accepted

I-vector has shown to be very effective in speaker verification with long-duration speech utterances. But when test utterances are of short duration, content mismatch between the enrollment and test utterances limit the performance of i-vector system. This paper proposes to extract local session var…

Cited by 0SourceScholar
2016

Exemplar-inspired strategies for low-resource spoken keyword search in Swahili

ICASSP 2016accepted

We present exemplar-inspired low-resource spoken keyword search strategies for acoustic modeling, keyword verification, and system combination. This state-of-the-art system was developed by the SINGA team in the context of the 2015 NIST Open Keyword Search Evaluation (OpenKWS15) using conversational…

Cited by 0SourceScholar
2016

Keyword search using query expansion for graph-based rescoring of hypothesized detections

ICASSP 2016accepted

In this work, we propose a novel framework for rescoring keyword search (KWS) detections using acoustic samples extracted from the training data. We view the keyword rescoring task as an information retrieval task and adopt the idea of query expansion. We expand a textual keyword with multiple speec…

Cited by 0SourceScholar
2016

Spoofing detection from a feature representation perspective

ICASSP 2016accepted

Spoofing detection, which discriminates the spoofed speech from the natural speech, has gained much attention recently. Low-dimensional features that are used in speaker recognition/verification are also used in spoofing detection. Unfortunately, they don't capture sufficient information required fo…

Cited by 0SourceScholar