← Search

Mirco Ravanelli

23 accepted papers

2026

FOCALCODEC-STREAM: STREAMING LOW-BITRATE SPEECH CODING VIA CAUSAL DISTILLATION

ICASSP 2026poster

Neural audio codecs are a fundamental component of modern generative audio pipelines. Although recent codecs achieve strong low-bitrate reconstruction and provide powerful representations for downstream tasks, most are non-streamable, limiting their use in real-time applications. We present FocalCod…

Cited by 0SourcePDFScholar
2026

Toward Faithful Explanations in Acoustic Anomaly Detection

ICASSP 2026poster

Interpretability is essential for user trust in real-world anomaly detection applications. However, deep learning models, despite their strong performance, often lack transparency. In this work, we study the interpretability of autoencoder-based models for audio anomaly detection, by comparing a sta…

Cited by 0SourcePDFScholar
2025

FocalCodec: Low-Bitrate Speech Coding via Focal Modulation Networks

NeurIPS 2025poster

Large language models have revolutionized natural language processing through self-supervised pretraining on massive datasets. Inspired by this success, researchers have explored adapting these methods to speech by discretizing continuous audio into tokens using neural audio codecs. However, existin…

Cited by 0SourcecodeScholar
2025

LMAC-TD: Producing Time Domain Explanations for Audio Classifiers

ICASSP 2025accepted

Neural networks are typically black-boxes that remain opaque with regards to their decision mechanisms. Several works in the literature have proposed post-hoc explanation methods to alleviate this issue. This paper proposes LMAC-TD, a post-hoc explanation method that trains a decoder to produce expl…

Cited by 0SourceScholar
2025

What Are They Doing? Joint Audio-Speech Co-Reasoning

ICASSP 2025accepted

In audio and speech processing, tasks usually focus on either the audio or speech modality, even when both sounds and human speech are present in the same audio clip. Recent Auditory Large Language Models (ALLMs) have made it possible to process audio and speech simultaneously within a single model,…

Cited by 0SourceScholar
2024

Adaptation Odyssey in LLMs: Why Does Additional Pretraining Sometimes Fail to Improve?

EMNLP 2024main

In the last decade, the generalization and adaptation abilities of deep learning models were typically evaluated on fixed training and test distributions. Contrary to traditional deep learning, large language models (LLMs) are (i) even more overparameterized, (ii) trained on unlabeled text corpora c…

Cited by 1SourcePDFScholar
2024

Listenable Maps for Zero-Shot Audio Classifiers

NeurIPS 2024poster

Interpreting the decisions of deep learning models, including audio classifiers, is crucial for ensuring the transparency and trustworthiness of this technology. In this paper, we introduce LMAC-ZS (Listenable Maps for Zero-Shot Audio Classifiers), which, to the best of our knowledge, is the first d…

Cited by 4SourcePDFScholar
2024

Resource-Efficient Separation Transformer

ICASSP 2024accepted

Transformers have recently achieved state-of-the-art performance in speech separation. These models, however, are computationally demanding and require a lot of learnable parameters. This paper explores Transformer-based speech separation with a reduced computational cost. Our main contribution is t…

Cited by 0SourceScholar
2024

TARIC-SLU: A Tunisian Benchmark Dataset for Spoken Language Understanding

COLING 2024main

In recent years, there has been a significant increase in interest in developing Spoken Language Understanding (SLU) systems. SLU involves extracting a list of semantic information from the speech signal. A major issue for SLU systems is the lack of sufficient amount of bi-modal (audio and textual s…

2024

Towards Foundational Models for Molecular Learning on Large-Scale Multi-Task Datasets

ICLR 2024poster

Recently, pre-trained foundation models have enabled significant advancements in multiple fields. In molecular machine learning, however, where datasets are often hand-curated, and hence typically small, the lack of datasets with labeled features, and codebases to manage those datasets, has hindered…

2023

Simulated Annealing in Early Layers Leads to Better Generalization

CVPR 2023poster

Recently, a number of iterative learning methods have been introduced to improve generalization. These typically rely on training for longer periods of time in exchange for improved generalization. LLF (later-layer-forgetting) is a state-of-the-art method in this category. It strengthens learning in…

2022

MetricGAN-U: Unsupervised Speech Enhancement/ Dereverberation Based Only on Noisy/ Reverberated Speech

ICASSP 2022accepted

Most of the deep learning-based speech enhancement models are learned in a supervised manner, which implies that pairs of noisy and clean speech are required during training. Consequently, several noisy speeches recorded in daily life cannot be used to train the model. Although certain unsupervised…

Cited by 0SourceScholar
2021

Attention Is All You Need In Speech Separation

ICASSP 2021accepted

Recurrent Neural Networks (RNNs) have long been the dominant architecture in sequence-to-sequence learning. RNNs, however, are inherently sequential models that do not allow parallelization of their computations. Transformers are emerging as a natural alternative to standard RNNs, replacing recurren…

Cited by 0SourceScholar
2021

Timers and Such: A Practical Benchmark for Spoken Language Understanding with Numbers

NeurIPS 2021poster

This paper introduces Timers and Such, a new open source dataset of spoken English commands for common voice control use cases involving numbers. We describe the gap in existing spoken language understanding datasets that Timers and Such fills, the design and creation of the dataset, and experiments…

Cited by 12SourcecodeScholar
2020

Multi-Task Self-Supervised Learning for Robust Speech Recognition

ICASSP 2020accepted

Despite the growing interest in unsupervised learning, extracting meaningful knowledge from unlabelled audio remains an open challenge. To take a step in this direction, we recently proposed a problem-agnostic speech encoder (PASE), that combines a convolutional encoder followed by multiple neural n…

Cited by 0SourceScholar
2020

Using Speech Synthesis to Train End-To-End Spoken Language Understanding Models

ICASSP 2020accepted

End-to-end models are an attractive new approach to spoken language understanding (SLU) in which the meaning of an utterance is inferred directly from the raw audio, without employing the standard pipeline composed of a separately trained speech recognizer and natural language understanding module.…

Cited by 0SourceScholar
2019

Quaternion Recurrent Neural Networks

ICLR 2019poster

Recurrent neural networks (RNNs) are powerful architectures to model sequential data, due to their capability to learn short and long-term dependencies between the basic elements of a sequence. Nonetheless, popular tasks such as speech or images recognition, involve multi-dimensional input features…

Cited by 183SourcePDFScholar
2017

A network of deep neural networks for Distant Speech Recognition

ICASSP 2017accepted

Despite the remarkable progress recently made in distant speech recognition, state-of-the-art technology still suffers from a lack of robustness, especially when adverse acoustic conditions characterized by non-stationary noises and reverberation are met. A prominent limitation of current systems li…

Cited by 0SourceScholar
2015

A multi-channel corpus for distant-speech interaction in presence of known interferences

ICASSP 2015accepted

This paper describes a new corpus of multi-channel audio data designed to study and develop distant-speech recognition systems able to cope with known interfering sounds propagating in an environment. The corpus consists of both real and simulated signals and of a corresponding detailed annotation.…

Cited by 0SourceScholar