← Search

Zehai Tu

7 accepted papers

2025

Can We "Cherry-Pick"? Investigating Multiple Renditions from a Generative Speech Synthesis Model

ICASSP 2025accepted

Generative Speech Models (GSMs) have seen a surge in popularity due to their ability to generate diverse and high-quality speech. Evaluating models that generate many different renditions for a given input sentence presents a new challenge. Listening tests are still the gold standard for evaluating…

Cited by 0SourceScholar
2025

Enabling Beam Search for Language Model-Based Text-to-Speech Synthesis

ICASSP 2025accepted

Tokenising continuous speech into sequences of discrete tokens and modelling them with language models (LMs) has led to significant success in text-to-speech (TTS) synthesis. Despite these models can generate speech with high quality and naturalness, their synthesised samples can still suffer from a…

Cited by 0SourceScholar
2025

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model

ICASSP 2025accepted

This paper introduces PFlow-VC, a conditional flow matching voice conversion model that leverages fine-grained discrete pitch tokens and target speaker prompt information for expressive voice conversion (VC). Previous VC works primarily focus on speaker conversion, with further exploration needed in…

Cited by 11SourceScholar
2023

The 2nd Clarity Enhancement Challenge for Hearing Aid Speech Intelligibility Enhancement: Overview and Outcomes

ICASSP 2023accepted

This paper reports on the design and outcomes of the 2nd Clarity Enhancement Challenge (CEC2), a challenge for stimulating novel approaches to hearing-aid speech intelligibility enhancement. The challenge was for a listener attending to a target speaker in a noisy, domestic environment. The challeng…

Cited by 0SourceScholar
2022

Auditory-Based Data Augmentation for end-to-end Automatic Speech Recognition

ICASSP 2022accepted

End-to-end models have achieved significant improvement on automatic speech recognition. One common method to improve performance of these models is expanding the data-space through data augmentation. Meanwhile, human auditory inspired front-ends have also demonstrated improvement for automatic spee…

Cited by 0SourceScholar