← Search

Robin San Roman

4 accepted papers

2025

Latent Watermarking of Audio Generative Models

ICASSP 2025accepted

The advancements in audio generative models have opened up new challenges in their responsible disclosure and the detection of their misuse. To address this, watermarking techniques have been recently developed, enabling the detection of content generated by a deployed model. For such techniques to…

Cited by 0SourceScholar
2025

MusicGen-Stem: Multi-stem music generation and edition through autoregressive modeling

ICASSP 2025accepted

While most music generation models generate a mixture of stems (in mono or stereo), we propose to train a multi-stem generative model with 3 stems (bass, drums and other) that learn the musical dependencies between them. To do so, we train one specialized compression algorithm per stem to tokenize t…

Cited by 0SourceScholar
2024

Proactive Detection of Voice Cloning with Localized Watermarking

ICML 2024poster

In the rapidly evolving field of speech generative models, there is a pressing need to ensure audio authenticity against the risks of voice cloning. We present AudioSeal, the first audio watermarking technique designed specifically for localized detection of AI-generated speech. AudioSeal employs a…

2023

From Discrete Tokens to High-Fidelity Audio Using Multi-Band Diffusion

NeurIPS 2023poster

Deep generative models can generate high-fidelity audio conditioned on various types of representations (e.g., mel-spectrograms, Mel-frequency Cepstral Coefficients (MFCC)). Recently, such models have been used to synthesize audio waveforms conditioned on highly compressed representations. Although…

Cited by 25SourcePDFScholar