← Search

Jing-Xuan Zhang

6 accepted papers

2026

STREAMING SPEECH RECOGNITION WITH DECODER-ONLY LARGE LANGUAGE MODELS AND LATENCY OPTIMIZATION

ICASSP 2026oral

Recent advances have demonstrated the potential of decoderonly large language models (LLMs) for automatic speech recognition (ASR). However, enabling streaming recognition within this framework remains a challenge. In this work, we propose a novel streaming ASR approach that integrates a read/write…

Cited by 0SourcePDFScholar
2023

Self-Supervised Audio-Visual Speech Representations Learning by Multimodal Self-Distillation

ICASSP 2023accepted

In this work, we present a novel method, named AV2vec, for learning audio-visual speech representations by multimodal self-distillation. AV2vec has a student and a teacher module, in which the student performs a masked latent feature regression task using the multimodal target features generated onl…

Cited by 0SourceScholar
2021

TaLNet: Voice Reconstruction from Tongue and Lip Articulation with Transfer Learning from Text-to-Speech Synthesis

AAAI 2021technical

This paper presents TaLNet, a model for voice reconstruction with ultrasound tongue and optical lip videos as inputs. TaLNet is based on an encoder-decoder architecture. Separate encoders are dedicated to processing the tongue and lip data streams respectively. The decoder pre…

Cited by 18SourcePDFScholar
2019

Dnn-based Spectral Enhancement for Neural Waveform Generators with Low-bit Quantization

ICASSP 2019accepted

This paper presents a spectral enhancement method to improve the quality of speech reconstructed by neural waveform generators with low-bit quantization. At training stage, this method builds a multiple-target DNN, which predicts log amplitude spectra of natural high-bit waveforms together with the…

Cited by 0SourceScholar
2019

Improving Sequence-to-sequence Voice Conversion by Adding Text-supervision

ICASSP 2019accepted

This paper presents methods of making using of text supervision to improve the performance of sequence-to-sequence (seq2seq) voice conversion. Compared with conventional frame-to-frame voice conversion approaches, the seq2seq acoustic modeling method proposed in our previous work achieved higher nat…

Cited by 0SourceScholar
2018

Forward Attention in Sequence- To-Sequence Acoustic Modeling for Speech Synthesis

ICASSP 2018accepted

This paper proposes a forward attention method for the sequence-to-sequence acoustic modeling of speech synthesis. This method is motivated by the nature of the monotonic alignment from phone sequences to acoustic sequences. Only the alignment paths that satisfy the monotonic condition are taken int…

Cited by 0SourceScholar