← Search

Florian Lux

5 accepted papers

2026

HOW TO LABEL RESYNTHESIZED AUDIO: THE DUAL ROLE OF NEURAL AUDIO CODECS IN AUDIO DEEPFAKE DETECTION

ICASSP 2026poster

Since Text-to-Speech systems typically don't produce waveforms directly, recent spoof detection studies use resynthesized waveforms from vocoders and neural audio codecs to simulate an attacker. Unlike vocoders, which are specifically designed for speech synthesis, neural audio codecs were originall…

Cited by 0SourcePDFScholar
2025

High-Resolution Speech Restoration with Latent Diffusion Model

ICASSP 2025accepted

Traditional speech enhancement methods often oversimplify the task of restoration by focusing on a single type of distortion. Generative models that handle multiple distortions frequently struggle with phone reconstruction and high-frequency harmonics, leading to breathing and gasping artifacts that…

Cited by 0SourceScholar
2023

Prosody Is Not Identity: A Speaker Anonymization Approach Using Prosody Cloning

ICASSP 2023accepted

Prosody is closely linked to the identity of a speaker, leading to individual pitch and intonation patterns. Therefore, it is challenging in speaker anonymization to generate speech utterances that both keep the original audio’s main prosodic structure and preserve the speaker’s privacy. In this pap…

Cited by 0SourceScholar
2022

Language-Agnostic Meta-Learning for Low-Resource Text-to-Speech with Articulatory Features

ACL 2022long

While neural text-to-speech systems perform remarkably well in high-resource scenarios, they cannot be applied to the majority of the over 6,000 spoken languages in the world due to a lack of appropriate training data. In this work, we use embeddings derived from articulatory vectors rather than emb…