← Search

Gaofeng Cheng

8 accepted papers

2025

Automatic Text Pronunciation Correlation Generation and Application for Contextual Biasing

ICASSP 2025accepted

Effectively distinguishing the pronunciation correlations between different written texts is a significant issue in linguistic acoustics. Traditionally, such pronunciation correlations are obtained through manually designed pronunciation lexicons. In this paper, we propose a data-driven method to au…

Cited by 0SourceScholar
2025

Hybrid Pseudo-Labeling for Semi-Supervised Automatic Speech Recognition

ICASSP 2025accepted

Pseudo-labeling based semi-supervised learning can mitigate the performance degradation resulting from the absence of labeled data in the target domain. In pseudo-labeling, the quality of pseudo-labels is crucial for the final performance. However, most works overlook the potential benefits of using…

Cited by 0SourceScholar
2025

SLIDE: Integrating Speech Language Model with LLM for Spontaneous Spoken Dialogue Generation

ICASSP 2025accepted

Recently, "textless" speech language models (SLMs) based on speech units have made huge progress in generating naturalistic speech, including non-verbal vocalizations. However, the generated speech samples often lack semantic coherence. In this paper, we propose SLM and LLM Integration for spontaneo…

Cited by 5SourceScholar
2022

Improving CTC-Based Speech Recognition Via Knowledge Transferring from Pre-Trained Language Models

ICASSP 2022accepted

Recently, end-to-end automatic speech recognition models based on connectionist temporal classification (CTC) have achieved impressive results, especially when fine-tuned from wav2vec2.0 models. Due to the conditional independence assumption, CTC-based models are always weaker than attention-based e…

Cited by 35SourceScholar
2022

Improving Non-Autoregressive End-to-End Speech Recognition with Pre-Trained Acoustic and Language Models

ICASSP 2022accepted

While Transformers have achieved promising results in end-to-end (E2E) automatic speech recognition (ASR), their autoregressive (AR) structure becomes a bottleneck for speeding up the decoding process. For real-world deployment, ASR systems are desired to be highly accurate while achieving fast infe…

Cited by 0SourceScholar
2021

History Utterance Embedding Transformer LM for Speech Recognition

ICASSP 2021accepted

History utterances contain rich contextual information; however, better extracting information from the history utterances and using it to improve the language model (LM) is still challenging. In this paper, we propose the history utterance embedding Transformer LM (HTLM), which includes an embeddin…

Cited by 0SourceScholar
2021

Pre-Training Transformer Decoder for End-to-End ASR Model with Unpaired Text Data

ICASSP 2021accepted

This paper presents a method to pre-train transformer-based encoder-decoder automatic speech recognition (ASR) models using sufficient target-domain text. During pre-training, we train the transformer decoder as a conditional language model with empty or artifical states, rather than the real encode…

Cited by 0SourceScholar
2020

Transformer-Based Online CTC/Attention End-To-End Speech Recognition Architecture

ICASSP 2020accepted

Recently, Transformer has gained success in automatic speech recognition (ASR) field. However, it is challenging to deploy a Transformer-based end-to-end (E2E) model for online speech recognition. In this paper, we propose the Transformer-based online CTC/attention E2E ASR architecture, which contai…

Cited by 0SourceScholar