← Search

Xingchen Song

6 accepted papers

2023

CB-Conformer: Contextual Biasing Conformer for Biased Word Recognition

ICASSP 2023accepted

Due to the mismatch between the source and target domains, how to better utilize the biased word information to improve the performance of the automatic speech recognition model in the target domain becomes a hot research topic. Previous approaches either decode with a fixed external language model…

Cited by 0SourceScholar
2023

Fast-U2++: Fast and Accurate End-to-End Speech Recognition in Joint CTC/Attention Frames

ICASSP 2023accepted

Recently, the unified streaming and non-streaming two-pass (U2/U2++) end-to-end model for speech recognition has shown great performance in terms of streaming capability, accuracy and latency. In this paper, we present fast-U2++, an enhanced version of U2++ to further reduce partial latency. The cor…

Cited by 0SourceScholar
2023

LightGrad: Lightweight Diffusion Probabilistic Model for Text-to-Speech

ICASSP 2023accepted

Recent advances in neural text-to-speech (TTS) models bring thousands of TTS applications into daily life, where models are deployed in cloud to provide services for customs. Among these models are diffusion probabilistic models (DPMs), which can be stably trained and are more parameter-efficient co…

Cited by 0SourceScholar
2023

TrimTail: Low-Latency Streaming ASR with Simple But Effective Spectrogram-Level Length Penalty

ICASSP 2023accepted

In this paper, we present TrimTail, a simple but effective emission regularization method to improve the latency of streaming ASR models. The core idea of TrimTail is to apply length penalty (i.e., by trimming trailing frames, see Fig. 1-(b)) directly on the spectrogram of input utterances, which do…

Cited by 0SourceScholar
2022

Pedestrian Intention Prediction Based on Traffic-Aware Scene Graph Model

IROS 2022poster

Anticipating the future behavior of pedestrians is a crucial part of deploying Automated Driving Systems (ADS) in urban traffic scenarios. Most recent works utilize a convolutional neural network (CNN) to extract visual information, which is then input to a recurrent neural network (RNN) along with…

Cited by 10SourceScholar
2021

Non-Autoregressive Transformer ASR with CTC-Enhanced Decoder Input

ICASSP 2021accepted

Non-autoregressive (NAR) transformer models have achieved significantly inference speedup but at the cost of inferior accuracy compared to autoregressive (AR) models in automatic speech recognition (ASR). Most of the NAR transformers take a fixed-length sequence filled with MASK tokens or a redundan…

Cited by 0SourceScholar