← Search

Peng Shen

4 accepted papers

2024

Hierarchical Cross-Modality Knowledge Transfer with Sinkhorn Attention for CTC-Based ASR

ICASSP 2024accepted

Due to the modality discrepancy between textual and acoustic modeling, efficiently transferring linguistic knowledge from a pretrained language model (PLM) to acoustic encoding for automatic speech recognition (ASR) still remains a challenging task. In this study, we propose a cross-modality knowled…

Cited by 0SourceScholar
2021

Unsupervised Neural Adaptation Model Based on Optimal Transport for Spoken Language Identification

ICASSP 2021accepted

Due to the mismatch of statistical distributions of acoustic speech between training and testing sets, the performance of spoken language identification (SLID) could be drastically degraded. In this paper, we propose an unsupervised neural adaptation model to deal with the distribution mismatch prob…

Cited by 0SourceScholar
2019

Interactive Learning of Teacher-student Model for Short Utterance Spoken Language Identification

ICASSP 2019accepted

Short utterance-based spoken language identification (LID) is a challenging task due to the large variation of its feature representation. Improving feature representation of short utterances using a teacher-student method has been shown its effectiveness for LID tasks. However, conventional teacher…

Cited by 0SourceScholar
2016

Local fisher discriminant analysis for spoken language identification

ICASSP 2016accepted

I-vector is a state-of-the-art technique widely used in spoken language identification systems. Since i-vectors include total variability factors, discriminant analysis methods have been introduced to find the most discriminative features while removing the undesired variables for language identific…

Cited by 0SourceScholar