← Search

Marco Tagliasacchi

11 accepted papers

2025

MAD Speech: Measures of Acoustic Diversity of Speech

NAACL 2025long

Generative spoken language models produce speech in a wide range of voices, prosody, and recording conditions, seemingly approaching the diversity of natural speech. However, the extent to which generated speech is acoustically diverse remains unclear due to a lack of appropriate metrics. We address…

2023

Disentangling Speech from Surroundings with Neural Embeddings

ICASSP 2023accepted

We present a method to separate speech signals from noisy environments in the embedding space of a neural audio codec. We introduce a new training procedure that allows our model to produce structured encodings of audio waveforms given by embedding vectors, where one part of the embedding vector rep…

Cited by 19SourceScholar
2023

LMCodec: A Low Bitrate Speech Codec with Causal Transformer Models

ICASSP 2023accepted

We introduce LMCodec, a causal neural speech codec that provides high quality audio at very low bitrates. The backbone of the system is a causal convolutional codec that encodes audio into a hierarchy of coarse-to-fine tokens using residual vector quantization. LMCodec trains a Transformer language…

Cited by 0SourceScholar
2021

LEAF: A Learnable Frontend for Audio Classification

ICLR 2021poster

Mel-filterbanks are fixed, engineered audio features which emulate human perception and have been used through the history of audio understanding up to today. However, their undeniable qualities are counterbalanced by the fundamental limitations of handmade representations. In this work we show that…

2021

Real-Time Speech Frequency Bandwidth Extension

ICASSP 2021accepted

In this paper we propose a lightweight model for frequency bandwidth extension of speech signals, increasing the sampling frequency from 8kHz to 16kHz while restoring the high frequency content to a level almost indistinguishable from the 16kHz ground truth. The model architecture is based on SEANet…

Cited by 0SourceScholar
2020

Pitch Estimation Via Self-Supervision

ICASSP 2020accepted

We present a method to estimate the fundamental frequency in monophonic audio, often referred to as pitch estimation. In contrast to existing methods, our neural network can be fully trained only on unlabeled data, using self-supervision. A tiny amount of labeled data is needed solely for mapping th…

Cited by 0SourceScholar
2016

Fast keypoint detection in video sequences

ICASSP 2016accepted

Several computer vision tasks exploit a succinct representation of the visual content in the form of sets of local features. Given an input image, feature extraction algorithms identify keypoints and assign to each of them a descriptor, based on the characteristics of the surrounding visual content.…

Cited by 0SourceScholar