Mapache: Masked Parallel Transformer for Advanced Speech Editing and Synthesis
Guillermo Cámbara, Patrick Lumban Tobing, Mikolaj Babianski, Ravichander Vipperla, Duo Wang, Ron Shmelkin, Giuseppe Coccia, Orazio Angelini
Abstract
Recent advancements in Generative AI, such as scaled Transformer large language models (LLM) and diffusion decoders, have revolutionized speech synthesis. With speech encompassing the complexities of natural language and audio dimensionality, many recent models have relied on autoregressive modeling of quantized speech tokens. Such an approach limits speech synthesis to left-to-right generation, making these models unsuitable for speech edits free from audio discontinuities. We introduce Mapache, a novel architecture that combines a non-autoregressive masked speech language model with acoustic diffusion modeling, offering a unique, fully parallel pipeline. Mapache excels in precise speech editing that is indiscernible to human listeners, exhibiting inpainting and zero-shot synthesis capabilities that either surpass or rival those of other state-of-the-art models that specialize in just one of these tasks. This paper also sheds light on optimizing the decoding process for such non-autoregressive models.
BibTeX
@inproceedings{icassp2024_mapachemaskedpar,
title = {Mapache: Masked Parallel Transformer for Advanced Speech Editing and Synthesis},
author = {Guillermo Cámbara and Patrick Lumban Tobing and Mikolaj Babianski and Ravichander Vipperla and Duo Wang and Ron Shmelkin and Giuseppe Coccia and Orazio Angelini and Arnaud Joly and Mateusz Lajszczak and Vincent Pollet},
booktitle = {ICASSP 2024},
year = {2024}
}