← Search

Haodong Zhou

3 accepted papers

2024

Reflow-TTS: A Rectified Flow Model for High-Fidelity Text-to-Speech

ICASSP 2024accepted

The diffusion models including Denoising Diffusion Probabilistic Models (DDPM) and score-based generative models have demonstrated excellent performance in speech synthesis tasks. However, its effectiveness comes at the cost of numerous sampling steps, resulting in prolonged sampling time required t…

Cited by 0SourceScholar
2023

Community Detection Graph Convolutional Network for Overlap-Aware Speaker Diarization

ICASSP 2023accepted

The clustering algorithm plays a crucial role in speaker diarization systems. However, traditional clustering algorithms suffer from the complex distribution of speaker embeddings and lack of digging potential relationships between speakers in a session. We propose a novel graph-based clustering app…

Cited by 0SourceScholar
2023

The XMU System for Audio-Visual Diarization and Recognition in MISP Challenge 2022

ICASSP 2023accepted

In this paper, we present our work in track 2 of the Multi-modal Information based Speech Processing (MISP) 2022 Challenge. We built a cascaded system and explored different acoustic front-ends and end-to-end speech recognition back-ends based on multimodal. To promote effective fusion between the d…

Cited by 0SourceScholar