← Search

Zhihao Du

13 accepted papers

2026

SYNCSPEECH: EFFICIENT AND LOW-LATENCY TEXT-TO-SPEECH BASED ON TEMPORAL MASKED TRANSFORMER

ICASSP 2026oral

Current text-to-speech (TTS) models face a persistent limitation: autoregressive (AR) models suffer from low generation efficiency, while modern non-autoregressive (NAR) models experience high latency due to their unordered temporal nature. To bridge this divide, we introduce SyncSpeech, an efficien…

Cited by 0SourcePDFScholar
2025

Build LLM-Based Zero-Shot Streaming TTS System with Cosyvoice

ICASSP 2025accepted

LLM-based text-to-speech(TTS) system has becoming the new trend and SOTA due to its high naturalness and zero-shot capability. However, it relies heavily on training data, usually requires at least thousands hours of labeled audio. In this report, we describe how to use pretrained CosyVoice model, t…

Cited by 0SourceScholar
2025

Enhancing Low-Resource ASR through Versatile TTS: Bridging the Data Gap

ICASSP 2025accepted

While automatic speech recognition (ASR) systems have achieved remarkable performance with large-scale datasets, their efficacy remains inadequate in low-resource settings, encompassing dialects, accents, minority languages, and long-tail hotwords, domains with significant practical relevance. With…

Cited by 0SourceScholar
2025

OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation

ACL 2025long

Full-duplex spoken dialogue systems significantly surpass traditional turn-based dialogue systems, as they allow simultaneous bidirectional communication, closely mirroring human-human interactions. However, achieving low latency and natural interactions in full-duplex dialogue systems remains a sig…

2025

Speech Recognition Meets Large Language Model: Benchmarking, Models, and Exploration

AAAI 2025technical

In this paper, we focus on prompting one of the most important tasks in the field of speech processing, i.e., automatic speech recognition (ASR), with speech foundation encoders and large language models (LLM). Despite the growing body of research in this area, we find that many crucial design decis…

2025

UniSpeaker: A Unified Approach for Multimodality-driven Speaker Generation

EMNLP 2025

While recent advances in reference-based speaker cloning have significantly improved the authenticity of synthetic speech, speaker generation driven by multimodal cues such as visual appearance, textual descriptions, and other biometric signals remains in its early stages. To pioneer truly multimoda

2024

FunCodec: A Fundamental, Reproducible and Integrable Open-Source Toolkit for Neural Speech Codec

ICASSP 2024accepted

This paper presents FunCodec, a fundamental neural speech codec toolkit, which is an extension of the open-source speech processing toolkit FunASR. FunCodec provides reproducible training recipes and inference scripts for the latest neural speech codec models, such as SoundStream and Encodec. Thanks…

Cited by 0SourceScholar
2022

M2Met: The Icassp 2022 Multi-Channel Multi-Party Meeting Transcription Challenge

ICASSP 2022accepted

Recent development of speech signal processing, such as speech recognition, speaker diarization, etc., has inspired numerous applications of speech technologies. The meeting scenario is one of the most valuable and, at the same time, most challenging scenarios for the deployment of speech technologi…

Cited by 0SourceScholar
2022

Speaker Overlap-aware Neural Diarization for Multi-party Meeting Analysis

EMNLP 2022main

Recently, hybrid systems of clustering and neural diarization models have been successfully applied in multi-party meeting analysis. However, current models always treat overlapped speaker diarization as a multi-label classification problem, where speaker dependency and overlaps are not well conside…

2022

Summary on the ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Grand Challenge

ICASSP 2022accepted

The ICASSP 2022 Multi-channel Multi-party Meeting Transcription Grand Challenge (M2MeT) focuses on one of the most valuable and the most challenging scenarios of speech technologies. The M2MeT challenge has particularly set up two tracks, speaker diarization (track 1) and multi-speaker automatic spe…

Cited by 0SourceScholar
2021

Capturing Temporal Dependencies Through Future Prediction for CNN-Based Audio Classifiers

ICASSP 2021accepted

This paper focuses on the problem of temporal dependency modeling in the CNN-based models for audio classification tasks. To capture audio temporal dependencies using CNNs, we take a different approach from the purely architecture-induced method and explicitly encode temporal dependencies into the C…

Cited by 2SourceScholar