← Search

Mingyu Cui

10 accepted papers

2026

MEASURING PROSODY DIVERSITY IN ZERO-SHOT TTS: A NEW METRIC, BENCHMARK, AND EXPLORATION

ICASSP 2026poster

Prosody diversity is essential for achieving naturalness and expressiveness in zero-shot text-to-speech (TTS). However, frequently used acoustic metrics capture only partial views of prosodic variation and correlate poorly with human perception, leaving the problem of reliably quantifying prosody di…

Cited by 10SourcePDFScholar
2025

Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC

ICASSP 2025accepted

Multi-talker speech recognition (MTASR) faces unique challenges in disentangling and transcribing overlapping speech. To address these challenges, this paper investigates the role of Connectionist Temporal Classification (CTC) in speaker disentanglement when incorporated with Serialized Output Train…

Cited by 0SourceScholar
2025

GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement

ACL 2025long

The evolution of speech technology has been spurred by the rapid increase in dataset sizes. Traditional speech models generally depend on a large amount of labeled training data, which is scarce for low-resource languages. This paper presents GigaSpeech 2, a large-scale, multi-domain, multilingual s…

2025

Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs

EMNLP 2025

With the rise of Speech Large Language Models (SpeechLLMs), two dominant approaches have emerged for speech processing: discrete tokens and continuous features. Each approach has demonstrated strong capabilities in audio-related processing tasks. However, the performance gap between these two paradi

Cited by 0SourcePDFScholar
2024

Cross-Speaker Encoding Network for Multi-Talker Speech Recognition

ICASSP 2024accepted

End-to-end multi-talker speech recognition has garnered great interest as an effective approach to directly transcribe overlapped speech from multiple speakers. Current methods typically adopt either 1) single-input multiple-output (SIMO) models with a branched encoder, or 2) single-input single-out…

Cited by 0SourceScholar
2024

Towards High-Performance and Low-Latency Feature-Based Speaker Adaptation of Conformer Speech Recognition Systems

ICASSP 2024accepted

Practical application of model-based speaker adaptation techniques to end-to-end ASR systems is hindered by speaker-level data scarcity and latency in speaker-dependent (SD) parameters update. To this end, data-efficient and low-latency rapid feature-based speaker adaptation approaches are proposed…

Cited by 0SourceScholar
2023

A Sidecar Separator Can Convert A Single-Talker Speech Recognition System to A Multi-Talker One

ICASSP 2023accepted

Although automatic speech recognition (ASR) can perform well in common non-overlapping environments, sustaining performance in multi-talker overlapping speech recognition remains challenging. Recent research revealed that ASR model’s encoder captures different levels of information with different la…

Cited by 0SourceScholar
2023

Exploring Self-Supervised Pre-Trained ASR Models for Dysarthric and Elderly Speech Recognition

ICASSP 2023accepted

Automatic recognition of disordered and elderly speech remains a highly challenging task to date due to the difficulty in collecting such data in large quantities. This paper explores a series of approaches to integrate domain adapted Self-Supervised Learning (SSL) pre-trained models into TDNN and C…

Cited by 0SourceScholar
2022

Exploiting Cross Domain Acoustic-to-Articulatory Inverted Features for Disordered Speech Recognition

ICASSP 2022accepted

Articulatory features are inherently invariant to acoustic signal distortion and have been successfully incorporated into automatic speech recognition (ASR) systems for normal speech. Their practical application to disordered speech recognition is often limited by the difficulty in collecting such s…

Cited by 0SourceScholar
2021

Neural Architecture Search for LF-MMI Trained Time Delay Neural Networks

ICASSP 2021accepted

Deep neural networks (DNNs) based automatic speech recognition (ASR) systems are often designed using expert knowledge and empirical evaluation. In this paper, a range of neural architecture search (NAS) techniques are used to automatically learn two types of hyper-parameters of state-of-the-art fac…

Cited by 28SourceScholar