← Search

Francesco Paissan

9 accepted papers

2026

FLEXIO: FLEXIBLE SINGLE- AND MULTI-CHANNEL SPEECH SEPARATION AND ENHANCEMENT

ICASSP 2026oral

Speech separation and enhancement (SSE) has advanced remarkably and achieved promising results in controlled settings, such as a fixed number of speakers and a fixed array configuration. Towards a universal SSE system, single-channel systems have been extended to deal with a variable number of speak…

Cited by 0SourcePDFScholar
2025

FocalCodec: Low-Bitrate Speech Coding via Focal Modulation Networks

NeurIPS 2025poster

Large language models have revolutionized natural language processing through self-supervised pretraining on massive datasets. Inspired by this success, researchers have explored adapting these methods to speech by discretizing continuous audio into tokens using neural audio codecs. However, existin…

Cited by 0SourcecodeScholar
2025

LMAC-TD: Producing Time Domain Explanations for Audio Classifiers

ICASSP 2025accepted

Neural networks are typically black-boxes that remain opaque with regards to their decision mechanisms. Several works in the literature have proposed post-hoc explanation methods to alleviate this issue. This paper proposes LMAC-TD, a post-hoc explanation method that trains a decoder to produce expl…

Cited by 0SourceScholar
2024

Listenable Maps for Zero-Shot Audio Classifiers

NeurIPS 2024poster

Interpreting the decisions of deep learning models, including audio classifiers, is crucial for ensuring the transparency and trustworthiness of this technology. In this paper, we introduce LMAC-ZS (Listenable Maps for Zero-Shot Audio Classifiers), which, to the best of our knowledge, is the first d…

Cited by 4SourcePDFScholar
2022

Scalable Neural Architectures for End-to-End Environmental Sound Classification

ICASSP 2022accepted

Sound Event Detection (SED) is a complex task simulating human ability to recognize what is happening in the surrounding from auditory signals only. This technology is a crucial asset in many applications such as smart cities. Here, urban sounds can be detected and processed by embedded devices in a…

Cited by 0SourceScholar