← Search

Ya-Jun Hu

8 accepted papers

2026

GROUP RELATIVE POLICY OPTIMIZATION FOR TEXT-TO-SPEECH WITH LARGE LANGUAGE MODELS

ICASSP 2026oral

This paper proposes a GRPO-based approach to enhance the performance of large language model (LLM)-based text-to-speech (TTS) models by deriving rewards from an off-the-shelf automatic speech recognition (ASR) model. Compared to previous reinforcement learning methods for LLM-based TTS, our method r…

Cited by 0SourcePDFScholar
2025

Anchored Monotonic Alignment and Representation Substitution for Rare Spontaneous Behaviors in Spontaneous Speech Synthesis

ICASSP 2025accepted

Spontaneous behaviors in speech pose significant challenges for speech synthesis. Existing research has not adequately addressed these behaviors, with most studies relying on specially recorded datasets. In contrast, real-world data more accurately reflects the natural, spontaneous speaking styles i…

Cited by 0SourceScholar
2025

Self-supervised Prosody Learning at Phoneme-level with Momentum Contrast for Speech Synthesis

ICASSP 2025accepted

This paper investigates leveraging large-scale speech data to enhance prosodic modeling in speech synthesis, and introduces a model named SP2MC which achieves self-supervised prosody learning at phoneme-level with momentum contrast. This model incorporates dual convolutional encoders for speech and…

Cited by 0SourceScholar
2022

Improving Recognition-Synthesis Based any-to-one Voice Conversion with Cyclic Training

ICASSP 2022accepted

In recognition-synthesis based any-to-one voice conversion (VC), an automatic speech recognition (ASR) model is employed to extract content-related features and a synthesizer is built to predict the acoustic features of the target speaker from the content-related features of any source speakers at t…

Cited by 0SourceScholar
2022

Neural Grapheme-To-Phoneme Conversion with Pre-Trained Grapheme Models

ICASSP 2022accepted

Neural network models have achieved state-of-the-art performance on grapheme-to-phoneme (G2P) conversion. However, their performance relies on large-scale pronunciation dictionaries, which may not be available for a lot of languages. Inspired by the success of the pre-trained language model BERT, th…

Cited by 15SourceScholar
2017

Extracting structural spectral features using what-where auto-encoders for statistical parametric speech synthesis

ICASSP 2017accepted

This paper presents a method to extract structural spectral features from spectral envelopes using what-where autoencoders (WWAE) for statistical parametric speech synthesis (SPSS). A WWAE is constructed by concatenating a convolutional net for input encoding and a deconvolutional net for reconstruc…

Cited by 0SourceScholar
2016

Deep belief network-based post-filtering for statistical parametric speech synthesis

ICASSP 2016accepted

The speech synthesized by statistical parametric speech synthesis (SPSS) always sounds muffled. One important reason is that the generated spectral envelopes are over-smoothed and many detailed spectral structures in natural speech are lost. This paper presents a deep belief network (DBN)-based post…

Cited by 0SourceScholar
2016

Modeling spectral envelopes using deep conditional restricted Boltzmann machines for statistical parametric speech synthesis

ICASSP 2016accepted

This paper proposes a spectral modeling method using a deep conditional restricted Boltzmann machine (DCRBM) for statistical parametric speech synthesis. In this method, a DCRBM, which combines a deep neural network (DNN) with a conditional restricted Boltzmann machine (CRBM), is utilized to describ…

Cited by 0SourceScholar