← Search

Yongqiang Wang

16 accepted papers

2026

Data-Centric Lessons To Improve Speech-Language Pretraining

ICLR 2026poster

Spoken Question-Answering (SQA) is a core capability for useful and interactive artificial intelligence systems. Recently, several speech-language models (SpeechLMs) have been released with a specific focus on improving their SQA performance. However, a lack of controlled ablations of pretraining da…

Cited by 0SourceScholar
2024

Locally Differentially Private Decentralized Stochastic Bilevel Optimization with Guaranteed Convergence Accuracy

ICML 2024poster

Decentralized bilevel optimization based machine learning techniques are achieving remarkable success in a wide variety of domains. However, the intensive exchange of information (involving nested-loops of consensus or communication iterations) in existing decentralized bilevel optimization algorith…

Cited by 6SourcePDFScholar
2024

Multilingual and Fully Non-Autoregressive ASR with Large Language Model Fusion: A Comprehensive Study

ICASSP 2024accepted

In the era of large models, the autoregressive nature of decoding often results in latency serving as a significant bottleneck. We propose a non-autoregressive LM-fused ASR system that effectively leverages the parallelization capabilities of accelerator hardware. Our approach combines the Universal…

Cited by 19SourceScholar
2024

USM-SCD: Multilingual Speaker Change Detection Based on Large Pretrained Foundation Models

ICASSP 2024accepted

We introduce a multilingual speaker change detection model (USM-SCD) that can simultaneously detect speaker turns and perform ASR for 96 languages. This model is adapted from a speech foundation model trained on a large quantity of supervised and unsupervised data, demonstrating the utility of fine-…

Cited by 0SourceScholar
2023

Accelerating RNN-T Training and Inference Using CTC Guidance

ICASSP 2023accepted

We propose a novel method to accelerate training and inference process of recurrent neural network transducer (RNN-T) based on the guidance from a co-trained connectionist temporal classification (CTC) model. We made a key assumption that if an encoder embedding frame is classified as a blank frame…

Cited by 0SourceScholar
2021

Emformer: Efficient Memory Transformer Based Acoustic Model for Low Latency Streaming Speech Recognition

ICASSP 2021accepted

This paper proposes an efficient memory transformer Emformer for low latency streaming speech recognition. In Emformer, the long-range history context is distilled into an augmented memory bank to reduce self-attention’s computation complexity. A cache mechanism saves the computation for the key and…

Cited by 0SourceScholar
2021

Streaming Simultaneous Speech Translation with Augmented Memory Transformer

ICASSP 2021accepted

Transformer-based models have achieved state-of-the-art performance on speech translation tasks. However, the model architecture is not efficient enough for streaming scenarios since self-attention is computed over an entire input sequence and the computational cost grows quadratically with the leng…

Cited by 0SourceScholar
2021

Transformer in Action: A Comparative Study of Transformer-Based Acoustic Models for Large Scale Speech Recognition Applications

ICASSP 2021accepted

Transformer-based acoustic models have shown promising results very recently. In this paper, we summarize the application of transformer and its streamable variant, Emformer based acoustic model [1] for large scale speech recognition applications. We compare the transformer based acoustic models wit…

Cited by 0SourceScholar
2020

DEJA-VU: Double Feature Presentation and Iterated Loss in Deep Transformer Networks

ICASSP 2020accepted

Deep acoustic models typically receive features in the first layer of the network, and process increasingly abstract representations in the subsequent layers. Here, we propose to feed the input features at multiple depths in the acoustic model. As our motivation is to allow acoustic models to re-exa…

Cited by 0SourceScholar
2020

Training ASR Models By Generation of Contextual Information

ICASSP 2020accepted

Supervised ASR models have reached unprecedented levels of accuracy, thanks in part to ever-increasing amounts of labelled training data. However, in many applications and locales, only moderate amounts of data are available, which has led to a surge in semi- and weakly-supervised learning research.…

Cited by 0SourceScholar
2020

Transformer-Based Acoustic Modeling for Hybrid Speech Recognition

ICASSP 2020accepted

We propose and evaluate transformer-based acoustic models (AMs) for hybrid speech recognition. Several modeling choices are discussed in this work, including various positional embedding methods and an iterated loss to enable training deep transformers. We also present a preliminary study of using l…

Cited by 0SourceScholar
2019

End-to-end Contextual Speech Recognition Using Class Language Models and a Token Passing Decoder

ICASSP 2019accepted

End-to-end modeling (E2E) of automatic speech recognition (ASR) blends all the components of a traditional speech recognition system into a single, unified model. Although it simplifies the ASR systems, the unified model is hard to adapt when training and testing data mismatches. In this work, we fo…

Cited by 0SourceScholar
2018

Towards End-to-end Spoken Language Understanding

ICASSP 2018accepted

Spoken language understanding system is traditionally designed as a pipeline of a number of components. First, the audio signal is processed by an automatic speech recognizer for transcription or n-best hypotheses. With the recognition results, a natural language understanding system classifies the…

Cited by 0SourceScholar
2016

Investigations on speaker adaptation of LSTM RNN models for speech recognition

ICASSP 2016accepted

Recently Long Short-Term Memory (LSTM) Recurrent Neural Networks (RNN) acoustic models have demonstrated superior performance over deep neural networks (DNN) models in speech recognition and many other tasks. Although a lot of work have been reported on DNN model adaptation, very little has been don…

Cited by 0SourceScholar
2016

Simplifying long short-term memory acoustic models for fast training and decoding

ICASSP 2016accepted

On acoustic modeling, recurrent neural networks (RNNs) using Long Short-Term Memory (LSTM) units have recently been shown to outperform deep neural networks (DNNs) models. This paper focuses on resolving two challenges faced by LSTM models: high model complexity and poor decoding efficiency. Motivat…

Cited by 0SourceScholar
2015

Small-footprint high-performance deep neural network-based speech recognition using split-VQ

ICASSP 2015accepted

Due to a large number of parameters in deep neural networks (DNNs), it is challenging to design a small-footprint DNN-based speech recognition system while maintaining a high recognition performance. Even with a singular value matrix decomposition (SVD) method and scalar quantization, the DNN model…

Cited by 0SourceScholar