← Search

Ming Tu

10 accepted papers

2026

ParaS2S: Benchmarking and Aligning Spoken Language Models for Paralinguistic-aware Speech-to-Speech Interaction

ICLR 2026poster

Speech-to-Speech (S2S) models have shown promising dialogue capabilities, but their ability to handle paralinguistic cues—such as emotion, tone, and speaker attributes—and to respond appropriately in both content and style remains underexplored. Progress is further hindered by the scarcity of high-q…

Cited by 0SourceScholar
2023

Efficient Neural Music Generation

NeurIPS 2023poster

Recent progress in music generation has been remarkably advanced by the state-of-the-art MusicLM, which comprises a hierarchy of three LMs, respectively, for semantic, coarse acoustic, and fine acoustic modelings. Yet, sampling with the MusicLM requires processing through these LMs one by one to obt…

2023

Streaming Voice Conversion via Intermediate Bottleneck Features and Non-Streaming Teacher Guidance

ICASSP 2023accepted

Streaming voice conversion (VC) is the task of converting the voice of one person to another in real-time. Previous streaming VC methods use phonetic posteriorgrams (PPGs) extracted from automatic speech recognition (ASR) systems to represent speaker-independent information. However, PPGs lack the p…

Cited by 0SourceScholar
2022

Cloning One's Voice Using Very Limited Data in the Wild

ICASSP 2022accepted

With the increasing popularity of speech synthesis products, the industry has put forward more requirements for personalized speech synthesis: (1) How to use low-resource, easily accessible data to clone a person’s voice. (2) How to clone a person’s voice while controlling the style and prosody. To…

Cited by 0SourceScholar
2020

Speaker-Invariant Affective Representation Learning via Adversarial Training

ICASSP 2020accepted

Representation learning for speech emotion recognition is challenging due to labeled data sparsity issue and lack of gold-standard references. In addition, there is much variability from input speech signals, human subjective perception of the signals and emotion label ambiguity. In this paper, we p…

Cited by 0SourceScholar
2018

Simulating Dysarthric Speech for Training Data Augmentation in Clinical Speech Applications

ICASSP 2018accepted

Training machine learning algorithms for speech applications requires large, labeled training data sets. This is problematic for clinical applications where obtaining such data is prohibitively expensive because of privacy concerns or lack of access. As a result, clinical speech applications typical…

Cited by 0SourceScholar
2016

Ranking the parameters of deep neural networks using the fisher information

ICASSP 2016accepted

The large number of parameters in deep neural networks (DNNs) often makes them prohibitive for low-power devices, such as field-programmable gate arrays (FPGA). In this paper, we propose a method to determine the relative importance of all network parameters by measuring the amount of information th…

Cited by 0SourceScholar