← Search

Nicola Pia

4 accepted papers

2025

FlowMAC: Conditional Flow Matching for Audio Coding at Low Bit Rates

ICASSP 2025accepted

This paper introduces FlowMAC, a novel neural audio codec for high-quality general audio compression at low bit rates based on conditional flow matching (CFM). FlowMAC jointly learns a mel spectrogram encoder, quantizer and decoder. At inference time the decoder integrates a continuous normalizing f…

Cited by 6SourceScholar
2025

Low-Resource Text-to-Speech Synthesis Using Noise-Augmented Training of ForwardTacotron

ICASSP 2025accepted

In recent years, several text-to-speech systems have been proposed to synthesize natural speech in zero-shot, few-shot, and low-resource scenarios. However, these methods typically require training with data from many different speakers. The speech quality across the speaker set typically is diverse…

Cited by 0SourceScholar
2022

PostGAN: A GAN-Based Post-Processor to Enhance the Quality of Coded Speech

ICASSP 2022accepted

The quality of speech coded by transform coding is affected by various artefacts especially when bitrates to quantize the frequency components become too low. In order to mitigate these coding artefacts and enhance the quality of coded speech, a post-processor that relies on a-priori information tra…

Cited by 0SourceScholar
2021

StyleMelGAN: An Efficient High-Fidelity Adversarial Vocoder with Temporal Adaptive Normalization

ICASSP 2021accepted

In recent years, neural vocoders have surpassed classical speech generation approaches in naturalness and perceptual quality of the synthesized speech. Computationally heavy models like WaveNet and WaveGlow achieve best results, while lightweight GAN models, e.g. MelGAN and Parallel WaveGAN, remain…

Cited by 0SourceScholar