← Search

Matthew Baas

3 accepted papers

2025

MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model

ICASSP 2025accepted

Codec-based text-to-speech (TTS) models have shown impressive quality with zero-shot voice cloning abilities. However, they often struggle with more expressive references or complex text inputs. We present MARS6, a robust encoder-decoder transformer for rapid, expressive TTS. MARS6 is built on recen…

Cited by 0SourceScholar
2025

kNN-SVC: Robust Zero-Shot Singing Voice Conversion with Additive Synthesis and Concatenation Smoothness Optimization

ICASSP 2025accepted

Robustness is critical in zero-shot singing voice conversion (SVC). This paper introduces two novel methods to strengthen the robustness of the kNN-VC framework for SVC. First, kNN-VC’s core representation, WavLM, lacks harmonic emphasis, resulting in dull sounds and ringing artifacts. To address th…

Cited by 0SourceScholar
2022

A Comparison of Discrete and Soft Speech Units for Improved Voice Conversion

ICASSP 2022accepted

The goal of voice conversion is to transform source speech into a target voice, keeping the content unchanged. In this paper, we focus on self-supervised representation learning for voice conversion. Specifically, we compare discrete and soft speech units as input features. We find that discrete rep…

Cited by 0SourceScholar