← Search

Shaoshi Ling

5 accepted papers

2024

Adapting Large Language Model with Speech for Fully Formatted End-to-End Speech Recognition

ICASSP 2024accepted

Most end-to-end (E2E) speech recognition models are composed of encoder and decoder blocks that perform acoustic and language modeling functions. Pretrained large language models (LLMs) have the potential to improve the performance of E2E ASR. However, integrating a pretrained language model into an…

Cited by 0SourceScholar
2022

Improving Pseudo-Label Training For End-To-End Speech Recognition Using Gradient Mask

ICASSP 2022accepted

In the recent trend of semi-supervised speech recognition, both self-supervised representation learning and pseudo-labeling have shown promising results. In this paper, we propose a novel approach to combine their ideas for end-to-end speech recognition model. Without any extra loss function, we uti…

Cited by 0SourceScholar
2020

Deep Contextualized Acoustic Representations for Semi-Supervised Speech Recognition

ICASSP 2020accepted

We propose a novel approach to semi-supervised automatic speech recognition (ASR). We first exploit a large amount of unlabeled audio data via representation learning, where we reconstruct a temporal slice of filterbank features from past and future context frames. The resulting deep contextualized…

Cited by 0SourceScholar