← Search

Ji Xu

12 accepted papers

2022

Improving CTC-Based Speech Recognition Via Knowledge Transferring from Pre-Trained Language Models

ICASSP 2022accepted

Recently, end-to-end automatic speech recognition models based on connectionist temporal classification (CTC) have achieved impressive results, especially when fine-tuned from wav2vec2.0 models. Due to the conditional independence assumption, CTC-based models are always weaker than attention-based e…

Cited by 35SourceScholar
2021

Analysis of Sensing Spectral for Signal Recovery under a Generalized Linear Model

NeurIPS 2021poster

We consider a nonlinear inverse problem $\mathbf{y}= f(\mathbf{Ax})$, where observations $\mathbf{y} \in \mathbb{R}^m$ are the componentwise nonlinear transformation of $\mathbf{Ax} \in \mathbb{R}^m$, $\mathbf{x} \in \mathbb{R}^n$ is the signal of interest and $\mathbf{A}$ is a known linear mapping.…

Cited by 10SourcePDFScholar
2021

When does preconditioning help or hurt generalization?

ICLR 2021poster

While second order optimizers such as natural gradient descent (NGD) often speed up optimization, their effect on generalization has been called into question. This work presents a more nuanced view on how the \textit{implicit bias} of optimizers affects the comparison of generalization properties.…

Cited by 50SourcePDFScholar
2019

Multiple Temporal Scales Based Speaker Embeddings Learning for Text-dependent Speaker Recognition

ICASSP 2019accepted

To extract high speaker-sensitive embeddings from deep neural networks is still a challenge in the field of speaker recognition. This paper proposes a novel network that learns speaker embeddings from multiple temporal scales. This idea comes from the recent biological research that the human audito…

Cited by 0SourceScholar
2018

A Deep Neural Network Based Method of Source Localization in a Shallow Water Environment

ICASSP 2018accepted

This paper applies deep neural network (DNN) to source localization in a shallow water environment because of its powerful modeling capability and the little dependence on the prior knowledge of environmental parameters. The classical two-stage scheme is adopted, in which feature extraction and DNN…

Cited by 0SourceScholar