← Search

Linhao Dong

8 accepted papers

2024

SA-SOT: Speaker-Aware Serialized Output Training for Multi-Talker ASR

ICASSP 2024accepted

Multi-talker automatic speech recognition plays a crucial role in scenarios involving multi-party interactions, such as meetings and conversations. Due to its inherent complexity, this task has been receiving increasing attention. Notably, the serialized output training (SOT) stands out among variou…

Cited by 0SourceScholar
2023

CIF-PT: Bridging Speech and Text Representations for Spoken Language Understanding via Continuous Integrate-and-Fire Pre-Training

ACL 2023findings

Speech or text representation generated by pre-trained models contains modal-specific information that could be combined for benefiting spoken language understanding (SLU) tasks. In this work, we propose a novel pre-training paradigm termed Continuous Integrate-and-Fire Pre-Training (CIF-PT). It rel…

2022

Improving End-to-End Contextual Speech Recognition with Fine-Grained Contextual Knowledge Selection

ICASSP 2022accepted

Nowadays, most methods for end-to-end contextual speech recognition bias the recognition process towards contextual knowledge. Since all-neural contextual biasing methods rely on phrase-level contextual modeling and attention-based relevance modeling, they may suffer from the confusion between simil…

Cited by 59SourceScholar
2021

Cif-Based Collaborative Decoding for End-to-End Contextual Speech Recognition

ICASSP 2021accepted

End-to-end (E2E) models have achieved promising results on multiple speech recognition benchmarks, and shown the potential to become the mainstream. However, the unified structure and the E2E training hamper injecting context information into them for contextual biasing. Though contextual LAS (CLAS)…

Cited by 0SourceScholar
2019

Self-attention Aligner: A Latency-control End-to-end Model for ASR Using Self-attention Network and Chunk-hopping

ICASSP 2019accepted

Self-attention network, an attention-based feedforward neural network, has recently shown the potential to replace recurrent neural networks (RNNs) in a variety of NLP tasks. However, it is not clear if the self-attention network could be a good alternative of RNNs in automatic speech recognition (A…

Cited by 0SourceScholar
2016

Variational Bayesian image fusion based on combined sparse representations

ICASSP 2016accepted

Hyper-spectral image fusion has been a hot topic in medical imaging and remote sensing. This paper proposes a Bayesian fusion model which combines the panchromatic (PAN) image and the low spatial resolution hyper-spectral (HS) image under the same framework. Sparsity constraint is introduced as doub…

Cited by 0SourceScholar