← Search

Qirong Mao

8 accepted papers

2022

Efficient Monaural Speech Separation with Multiscale Time-Delay Sampling

ICASSP 2022accepted

Recently, the segmented sample-level modeling approach based on Dual-Path Recurrent Neural Network (DPRNN) has been proved to be effective in Monaural Speech Separation (MSS). Many dual-path networks such as Dual-Path Transformer Network (DPTNet), with a series of improvements to DPRNN, have also im…

Cited by 0SourceScholar
2022

Statistical Pyramid Dense Time Delay Neural Network for Speaker Verification

ICASSP 2022accepted

Recently, speaker verification (SV) techniques relay on deep learning frameworks to extract more informative embedding vectors, which greatly improves the accuracy compared with traditional machine learning methods. The well-known x-vector architecture, a time delay neural network (TDNN), is widely…

Cited by 0SourceScholar
2018

Joint Pose and Expression Modeling for Facial Expression Recognition

CVPR 2018poster

Facial expression recognition (FER) is a challenging task due to different expressions under arbitrary poses. Most conventional approaches either perform face frontalization on a non-frontal facial image or learn separate classifiers for each pose. Different from existing methods, in this paper, we…

2016

Domain adaptation for speech emotion recognition by sharing priors between related source and target classes

ICASSP 2016accepted

In speech emotion recognition (SER), speech data is usually captured from different scenarios, which often leads to significant performance degradation due to the inherent mismatch between training and test set. To cope with this problem, we propose a domain adaptation method called Sharing Priors b…

Cited by 0SourceScholar