← Search

Zijian Yang

6 accepted papers

2026

Sequence-Level Unsupervised Training in Speech Recognition: A Theoretical Study

ICASSP 2026poster

Unsupervised speech recognition is a task of training a speech recognition model with unpaired data. To determine when and how unsupervised speech recognition can succeed, and how classification error relates to candidate training objectives, we develop a theoretical framework for unsupervised speec…

Cited by 0SourcePDFScholar
2025

Classification Error Bound for Low Bayes Error Conditions in Machine Learning

ICASSP 2025accepted

In statistical classification and machine learning, classification error is an important performance measure, which is minimized by the Bayes decision rule. In practice, the unknown true distribution is usually replaced with a model distribution estimated from the training data in the Bayes decision…

Cited by 0SourceScholar
2025

Efficient Supernet Training with Orthogonal Softmax for Scalable ASR Model Compression

ICASSP 2025accepted

ASR systems are deployed across diverse environments, each with specific hardware constraints. We use supernet training to jointly train multiple encoders of varying sizes, enabling dynamic model size adjustment to fit hardware constraints without redundant training. Moreover, we introduce a novel m…

Cited by 0SourceScholar
2024

On the Relation Between Internal Language Model and Sequence Discriminative Training for Neural Transducers

ICASSP 2024accepted

Internal language model (ILM) subtraction has been widely applied to improve the performance of the RNN-Transducer with external language model (LM) fusion for speech recognition. In this work, we show that sequence discriminative training has a strong correlation with ILM subtraction from both theo…

Cited by 0SourceScholar
2023

Generalized Discriminative Deep Non-Negative Matrix Factorization Based on Latent Feature and Basis Learning

IJCAI 2023poster

As a powerful tool for data representation, deep NMF has attracted much attention in recent years. Current deep NMF builds the multi-layer structure by decomposing either basis matrix or feature matrix into multiple factors, and probably complicates the learning process when data is insufficient or…

2023

Lattice-Free Sequence Discriminative Training for Phoneme-Based Neural Transducers

ICASSP 2023accepted

Recently, RNN-Transducers have achieved remarkable results on various automatic speech recognition tasks. However, lattice-free sequence discriminative training methods, which obtain superior performance in hybrid models, are rarely investigated in RNN-Transducers. In this work, we propose three lat…

Cited by 0SourceScholar