← Search

Mingbo Ma

6 accepted papers

2026

Scaling Speech Tokenizers with Diffusion Autoencoders

ICLR 2026poster

Speech tokenizers are foundational to speech language models, yet existing approaches face two major challenges: (1) balancing trade-offs between encoding semantics for understanding and acoustics for reconstruction, and (2) achieving low bit rates and low token rates. We propose Speech Diffusion To…

Cited by 0SourceScholar
2025

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

ICLR 2025poster

The imitation of voice, targeted on specific speech attributes such as timbre and speaking style, is crucial in speech generation. However, existing methods rely heavily on annotated data, and struggle with effectively disentangling timbre and style, leading to challenges in achieving controllable g…

2023

Efficient Neural Music Generation

NeurIPS 2023poster

Recent progress in music generation has been remarkably advanced by the state-of-the-art MusicLM, which comprises a hierarchy of three LMs, respectively, for semantic, coarse acoustic, and fine acoustic modelings. Yet, sampling with the MusicLM requires processing through these LMs one by one to obt…

2022

A$^3$T: Alignment-Aware Acoustic and Text Pretraining for Speech Synthesis and Editing

ICML 2022spotlight

Recently, speech representation learning has improved many speech-related tasks such as speech recognition, speech classification, and speech-to-text translation. However, all the above tasks are in the direction of speech understanding, but for the inverse direction, speech synthesis, the potential…

2021

Fused Acoustic and Text Encoding for Multimodal Bilingual Pretraining and Speech Translation

ICML 2021spotlight

Recently, representation learning for text and speech has successfully improved many language related tasks. However, all existing methods suffer from two limitations: (a) they only learn from one input modality, while a unified representation for both speech and text is needed by tasks such as end-…

2021

Improving Simultaneous Translation by Incorporating Pseudo-References with Fewer Reorderings

EMNLP 2021main

Simultaneous translation is vastly different from full-sentence translation, in the sense that it starts translation before the source sentence ends, with only a few words delay. However, due to the lack of large-scale, high-quality simultaneous translation datasets, most such systems are still trai…

Cited by 21SourcePDFScholar