← Search

Li-Rong Dai

38 accepted papers

2025

Dynamic SRM Curriculum for Trustworthy Multi-modal Classification

ICASSP 2025accepted

Trustworthy multi-modal learning integrates multiple sources of data reliably. However, the current methods still focus on performance improvement by developing deep multi-modal networks. These approaches frequently encounter challenges due to the inherent non-convex nature of deep neural networks a…

Cited by 0SourceScholar
2025

Trusted Mamba Contrastive Network for Multi-View Clustering

ICASSP 2025accepted

Multi-view clustering can partition data samples into their categories by learning a consensus representation in an unsupervised way and has received more and more attention in recent years. However, there is an untrusted fusion problem. The reasons for this problem are as follows: 1) The current me…

Cited by 0SourceScholar
2024

Adaptive Confidence Multi-View Hashing for Multimedia Retrieval

ICASSP 2024accepted

The multi-view hash method converts heterogeneous data from multiple views into binary hash codes, which is one of the critical technologies in multimedia retrieval. However, the current methods mainly explore the complementarity among multiple views while lacking confidence in learning and fusion.…

Cited by 0SourceScholar
2023

A Multi-Scale Feature Aggregation Based Lightweight Network for Audio-Visual Speech Enhancement

ICASSP 2023accepted

Audio-visual speech enhancement (AVSE) was shown to be superior over conventional audio-only counterpart for improving the speech quality. However, most existing AVSE models are heavyweight in the sense of parameter count, which is inappropriate for the deployment and practical applications. In this…

Cited by 0SourceScholar
2023

AST-SED: An Effective Sound Event Detection Method Based on Audio Spectrogram Transformer

ICASSP 2023accepted

In this paper, we propose an effective sound event detection (SED) method based on the audio spectrogram transformer (AST) model, pretrained on the large-scale AudioSet for audio tagging (AT) task, termed AST-SED. Pretrained AST models have recently shown promise on DCASE2022 challenge task4 where t…

Cited by 0SourceScholar
2023

Joint Generative-Contrastive Representation Learning for Anomalous Sound Detection

ICASSP 2023accepted

In this paper, we propose a joint generative and contrastive representation learning method (GeCo) for anomalous sound detection (ASD). GeCo exploits a Predictive AutoEncoder (PAE) equipped with self-attention as a generative model to perform frame-level prediction. The output of the PAE together wi…

Cited by 0SourceScholar
2023

Robust Data2VEC: Noise-Robust Speech Representation Learning for ASR by Combining Regression and Improved Contrastive Learning

ICASSP 2023accepted

Self-supervised pre-training methods based on contrastive learning or regression tasks can utilize more unlabeled data to improve the performance of automatic speech recognition (ASR). However, the robustness impact of combining the two pre-training tasks and constructing different negative samples…

Cited by 0SourceScholar
2023

Stargan-vc Based Cross-Domain Data Augmentation for Speaker Verification

ICASSP 2023accepted

Automatic speaker verification (ASV) faces domain shift caused by the mismatch of intrinsic and extrinsic factors, such as recording device and speaking style, in real-world applications, which leads to severe performance degradation. Since single-speaker multi-condition (SSMC) data is difficult to…

Cited by 0SourceScholar
2022

A Noise-Robust Self-Supervised Pre-Training Model Based Speech Representation Learning for Automatic Speech Recognition

ICASSP 2022accepted

Wav2vec2.0 is a popular self-supervised pre-training framework for learning speech representations in the context of automatic speech recognition (ASR). It was shown that wav2vec2.0 has a good robustness against the domain shift, while the noise robustness is still unclear. In this work, we therefor…

Cited by 0SourceScholar
2022

Domain Robust Deep Embedding Learning for Speaker Recognition

ICASSP 2022accepted

This paper presents a domain robust deep embedding learning method for speaker verification (SV) tasks. Most recent methods utilize deep neural networks (DNN) to learn compact and discriminative speaker embeddings from large-scale labeled datasets such as VoxCeleb and the NIST SRE corpus. Despite th…

Cited by 0SourceScholar
2022

Frontend Attributes Disentanglement for Speech Emotion Recognition

ICASSP 2022accepted

Speech emotion recognition (SER) with limited size dataset is a challenging task, since a spoken utterance contains various disturbing attributes besides emotion, including speaker, content, and language. However, due to a close relationship between speaker and emotion attributes, simply fine-tuning…

Cited by 0SourceScholar
2022

Reference Microphone Selection and Low-Rank Approximation Based Multichannel Wiener Filter with Application to Speech Recognition

ICASSP 2022accepted

For multichannel speech recognition systems, it is necessary to use a speech enhancement module to suppress ambient noises. Given second-order statistics, the multichannel Wiener filter (MWF) can be designed for noise reduction. It was shown that the MWF noise reduction performance depends on the se…

Cited by 0SourceScholar
2022

Self-Supervised Representation Learning for Unsupervised Anomalous Sound Detection Under Domain Shift

ICASSP 2022accepted

In this paper, a self-supervised representation learning method is proposed for anomalous sound detection (ASD). ASD has received much research attention in recent DCASE challenges. It aims to identify whether a sound emitted from a machine is anomalous or not, given only normal sound data. This is…

Cited by 0SourceScholar
2022

Supervised and Self-Supervised Pretraining Based Covid-19 Detection Using Acoustic Breathing/Cough/Speech Signals

ICASSP 2022accepted

A rapid-accurate detection method for COVID-19 is rather important for avoiding its pandemic. In this work, we propose a bi-directional long short-term memory (BiLSTM) network based COVID-19 detection method using breath/speech/cough signals. Three kinds of acoustic signals are taken to train the ne…

Cited by 0SourceScholar
2021

An Effective Deep Embedding Learning Method Based on Dense-Residual Networks for Speaker Verification

ICASSP 2021accepted

In this paper, we present an effective end-to-end deep embedding learning method based on Dense-Residual networks, which combine the advantages of a densely connected convolutional network (DenseNet) and a residual network (ResNet), for speaker verification (SV). Unlike a model ensemble strategy whi…

Cited by 0SourceScholar
2021

An Improved Mean Teacher Based Method for Large Scale Weakly Labeled Semi-Supervised Sound Event Detection

ICASSP 2021accepted

This paper presents an improved mean teacher (MT) based method for large-scale weakly labeled semi-supervised sound event detection (SED), by focusing on learning a better student model. Two main improvements are proposed based on the authors’ previous perturbation based MT method. Firstly, an event…

Cited by 26SourceScholar
2020

An Online Speaker-aware Speech Separation Approach Based on Time-domain Representation

ICASSP 2020accepted

Despite the significant progress of deep learning based speech separation methods, it remains challenging to extract and track the speech from target speakers, especially in a single-channel multiple speaker situation. Previously, the authors proposed a source-aware context network to exploit the te…

Cited by 0SourceScholar
2020

Extracting Unit Embeddings Using Sequence-To-Sequence Acoustic Models for Unit Selection Speech Synthesis

ICASSP 2020accepted

This paper presents a method of using the intermediate representations between linguistic and acoustic features in a Tacotron model to derive the cost functions for unit selection speech synthesis. By extracting the outputs of the Tacotron encoder, each phone-sized candidate unit in the corpus is re…

Cited by 0SourceScholar
2020

Task-Aware Mean Teacher Method for Large Scale Weakly Labeled Semi-Supervised Sound Event Detection

ICASSP 2020accepted

Weakly labeled semi-supervised learning methods have recently drawn increasing attention from the research community for sound event detection tasks. Due to the weakness of the labelling, neural networks are often designed to perform sound event detection (SED) and audio tagging (AT) at the same tim…

Cited by 0SourceScholar
2019

A Region Based Attention Method for Weakly Supervised Sound Event Detection and Classification

ICASSP 2019accepted

Recently, an attention based convolutional recurrent neural network (CRNN) with learnable gated linear units (GLUs) has achieved state-of-the-art performance for audio tagging (AT) and sound event detection (SED) tasks in the Detection and Classification of Acoustic Scenes and Events (DCASE) challen…

Cited by 0SourceScholar
2019

Improving Sequence-to-sequence Voice Conversion by Adding Text-supervision

ICASSP 2019accepted

This paper presents methods of making using of text supervision to improve the performance of sequence-to-sequence (seq2seq) voice conversion. Compared with conventional frame-to-frame voice conversion approaches, the seq2seq acoustic modeling method proposed in our previous work achieved higher nat…

Cited by 0SourceScholar
2018

Densely Connected Progressive Learning for LSTM-Based Speech Enhancement

ICASSP 2018accepted

Recently, we proposed a novel progressive learning (PL) framework for deep neural network (DNN) based speech enhancement to improve the performance in low signal-to-noise ratio (SNR) environments. In this study, several new contributions are made to this framework. First, the advanced long short-ter…

Cited by 0SourceScholar
2018

Forward Attention in Sequence- To-Sequence Acoustic Modeling for Speech Synthesis

ICASSP 2018accepted

This paper proposes a forward attention method for the sequence-to-sequence acoustic modeling of speech synthesis. This method is motivated by the nature of the monotonic alignment from phone sequences to acoustic sequences. Only the alignment paths that satisfy the monotonic condition are taken int…

Cited by 0SourceScholar
2018

Source-Aware Context Network for Single-Channel Multi-Speaker Speech Separation

ICASSP 2018accepted

Deep learning based approaches have achieved promising performance in speaker-dependent single-channel multispeaker speech separation. However, partly due to the label permutation problem, they may encounter difficulties in speaker-independent conditions. Recent methods address this problem by some…

Cited by 0SourceScholar
2017

Adaptation of PLDA for multi-source text-independent speaker verification

ICASSP 2017accepted

Probabilistic linear discriminant analysis (PLDA) is widely described as an effective model for text-independent speaker verification in the i-vector space. The PLDA scoring function is typically formulated as the likelihood ratio between the speaker-adapted and the universal PLDAs. In this case, th…

Cited by 0SourceScholar
2017

Extracting structural spectral features using what-where auto-encoders for statistical parametric speech synthesis

ICASSP 2017accepted

This paper presents a method to extract structural spectral features from spectral envelopes using what-where autoencoders (WWAE) for statistical parametric speech synthesis (SPSS). A WWAE is constructed by concatenating a convolutional net for input encoding and a deconvolutional net for reconstruc…

Cited by 0SourceScholar
2016

Compact convolutional neural network transfer learning for small-scale image classification

ICASSP 2016accepted

Transfer learning methods have demonstrated state-of-the-art performance on various small-scale image classification tasks. This is generally achieved by exploiting the information from an ImageNet convolution neural network (ImageNet CNN). However, the transferred CNN model is generally with high c…

Cited by 0SourceScholar
2016

Content-aware local variability vector for speaker verification with short utterance

ICASSP 2016accepted

I-vector has shown to be very effective in speaker verification with long-duration speech utterances. But when test utterances are of short duration, content mismatch between the enrollment and test utterances limit the performance of i-vector system. This paper proposes to extract local session var…

Cited by 0SourceScholar
2016

Deep belief network-based post-filtering for statistical parametric speech synthesis

ICASSP 2016accepted

The speech synthesized by statistical parametric speech synthesis (SPSS) always sounds muffled. One important reason is that the generated spectral envelopes are over-smoothed and many detailed spectral structures in natural speech are lost. This paper presents a deep belief network (DBN)-based post…

Cited by 0SourceScholar
2016

Modeling spectral envelopes using deep conditional restricted Boltzmann machines for statistical parametric speech synthesis

ICASSP 2016accepted

This paper proposes a spectral modeling method using a deep conditional restricted Boltzmann machine (DCRBM) for statistical parametric speech synthesis. In this method, a DCRBM, which combines a deep neural network (DNN) with a conditional restricted Boltzmann machine (CRBM), is utilized to describ…

Cited by 0SourceScholar
2016

Modulation spectrum compensation for HMM-based speech synthesis using line spectral pairs

ICASSP 2016accepted

In previous work, a method to compensate the divergence between the distributions of natural and generated modulation spectra (MS) has been proposed for hidden Markov model (HMM) based speech synthesis. This method can alleviate the over-smoothing effect of parameter generation when Mel-cepstral coe…

Cited by 0SourceScholar
2016

Speaker adaptation OF RNN-BLSTM for speech recognition based on speaker code

ICASSP 2016accepted

Recently, recurrent neural network with bidirectional Long Short-Term Memory (RNN-BLSTM) acoustic model has been shown to give great performance on the TIMIT [1] and other speech recognition tasks. Meanwhile, the speaker code based adaptation method has been demonstrated as a valid adaptation method…

Cited by 0SourceScholar
2015

Channel adaptation of plda for text-independent speaker verification

ICASSP 2015accepted

Probabilistic linear discriminant analysis (PLDA) has shown to be effective for modeling channel variability in the i-vector space for text-independent speaker verification. Speaker verification is a binary hypothesis testing. Given a test segment, the verification score could be computed as the log…

Cited by 0SourceScholar
2015

Improved language identification using deep bottleneck network

ICASSP 2015accepted

Effective representation plays an important role in automatic spoken language identification (LID). Recently, several representations that employ a pre-trained deep neural network (DNN) as the front-end feature extractor, have achieved state-of-the-art performance. However the performance is still f…

Cited by 0SourceScholar
2015

Joint training of front-end and back-end deep neural networks for robust speech recognition

ICASSP 2015accepted

Based on the recently proposed speech pre-processing front-end with deep neural networks (DNNs), we first investigate different feature mapping directly from noisy speech via DNN for robust speech recognition. Next, we propose to jointly train a single DNN for both feature mapping and acoustic model…

Cited by 0SourceScholar
2015

Spectral conversion using deep neural networks trained with multi-source speakers

ICASSP 2015accepted

This paper presents a method for voice conversion using deep neural networks (DNNs) trained with multiple source speakers. The proposed DNNs can be used in two ways for different scenarios: 1) in the absence of training data for source speaker, the DNNs can be treated as source-speaker-independent m…

Cited by 0SourceScholar
2015

Speech Separation based on signal-noise-dependent deep neural networks for robust speech recognition

ICASSP 2015accepted

In this paper, we propose a new signal-noise-dependent (SND) deep neural network (DNN) framework to further improve the separation and recognition performance of the recently developed technique for general DNN-based speech separation. We adopt a divide and conquer strategy to design the proposed SN…

Cited by 0SourceScholar
2015

Unsupervised speaker adaptation of deep neural network based on the combination of speaker codes and singular value decomposition for speech recognition

ICASSP 2015accepted

Recently, we have proposed a general adaptation scheme for deep neural network based on discriminant condition codes and applied it to supervised speaker adaptation in speech recognition based on either frame-level cross-entropy or sequence-level maximum mutual information training criterion [1, 2,…

Cited by 0SourceScholar