← Search

Joan Serrà

23 accepted papers

2026

LEVERAGING WHISPER EMBEDDINGS FOR AUDIO-BASED LYRICS MATCHING

ICASSP 2026poster

Audio-based lyrics matching can be an appealing alternative to other content-based retrieval approaches, but existing methods often suffer from limited reproducibility and inconsistent baselines. In this work, we introduce WEALY, a fully reproducible pipeline that leverages Whisper decoder embedding…

Cited by 0SourcePDFScholar
2025

Joint Semantic Knowledge Distillation and Masked Acoustic Modeling for Full-band Speech Restoration With Improved Intelligibility

ICASSP 2025accepted

Speech restoration aims at restoring full-band speech with high quality and intelligibility, considering a diverse set of distortions. MaskSR is a recently proposed generative model for this task. As other models of its kind, MaskSR attains high quality but, as we show, intelligibility can be substa…

Cited by 0SourceScholar
2025

Supervised Contrastive Learning from Weakly-Labeled Audio Segments for Musical Version Matching

ICML 2025poster

Detecting musical versions (different renditions of the same piece) is a challenging task with important applications. Because of the ground truth nature, existing approaches match musical versions at the track level (e.g., whole song). However, most applications require to match them at the segment…

Cited by 0SourcePDFScholar
2024

GASS: Generalizing Audio Source Separation with Large-Scale Data

ICASSP 2024accepted

Universal source separation targets at separating the audio sources of an arbitrary mix, removing the constraint to operate on a specific domain like speech or music. Yet, the potential of universal source separation is limited because most existing works focus on mixes with predominantly sound even…

Cited by 0SourceScholar
2024

Masked Generative Video-to-Audio Transformers with Enhanced Synchronicity

ECCV 2024poster

"Video-to-audio (V2A) generation leverages visual-only video features to render plausible sounds that match the scene. Importantly, the generated sound onsets should match the visual actions that are aligned with them, otherwise unnatural synchronization artifacts arise. Recent works have explored t…

2023

Adversarial Permutation Invariant Training for Universal Sound Separation

ICASSP 2023accepted

Universal sound separation consists of separating mixes with arbitrary sounds of different types, and permutation invariant training (PIT) is used to train source agnostic models that do so. In this work, we complement PIT with adversarial losses but find it challenging with the standard formulation…

Cited by 0SourceScholar
2023

Full-Band General Audio Synthesis with Score-Based Diffusion

ICASSP 2023accepted

Recent works have shown the capability of deep generative models to tackle general audio synthesis from a single label, producing a variety of impulsive, tonal, and environmental sounds. Such models operate on band-limited signals and, as a result of an autoregressive approach, they are typically co…

Cited by 0SourceScholar
2023

Quantitative Evidence on Overlooked Aspects of Enrollment Speaker Embeddings for Target Speaker Separation

ICASSP 2023accepted

Single channel target speaker separation (TSS) aims at extracting a speaker’s voice from a mixture of multiple talkers given an enrollment utterance of that speaker. A typical deep learning TSS framework consists of an upstream model that obtains enrollment speaker embeddings and a downstream model…

Cited by 0SourceScholar
2022

On Loss Functions and Evaluation Metrics for Music Source Separation

ICASSP 2022accepted

We investigate which loss functions provide better separations via benchmarking an extensive set of those for music source separation. To that end, we first survey the most representative audio source separation losses we identified, to later consistently benchmark them in a controlled experimental…

Cited by 0SourceScholar
2021

Automatic Multitrack Mixing With A Differentiable Mixing Console Of Neural Audio Effects

ICASSP 2021accepted

Applications of deep learning to automatic multitrack mixing are largely unexplored. This is partly due to the limited available data, coupled with the fact that such data is relatively unstructured and variable. To address these challenges, we propose a domain-inspired model with a strong inductive…

Cited by 0SourceScholar
2021

Investigating the Efficacy of Music Version Retrieval Systems for Setlist Identification

ICASSP 2021accepted

The setlist identification (SLI) task addresses a music recognition use case where the goal is to retrieve the metadata and times-tamps for all the tracks played in live music events. Due to various musical and non-musical changes in live performances, developing automatic SLI systems is still a cha…

Cited by 0SourceScholar
2020

Accurate and Scalable Version Identification Using Musically-Motivated Embeddings

ICASSP 2020accepted

The version identification (VI) task deals with the automatic detection of recordings that correspond to the same underlying musical piece. Despite many efforts, VI is still an open problem, with much room for improvement, specially with regard to combining accuracy and scalability. In this paper, w…

Cited by 0SourceScholar
2020

Input Complexity and Out-of-distribution Detection with Likelihood-based Generative Models

ICLR 2020poster

Likelihood-based generative models are a promising resource to detect out-of-distribution (OOD) inputs which could compromise the robustness or reliability of a machine learning system. However, likelihoods derived from such models have been shown to be problematic for detecting certain types of inp…

Cited by 315SourceScholar
2019

Blow: a single-scale hyperconditioned flow for non-parallel raw-audio voice conversion

NeurIPS 2019poster

End-to-end models for raw audio generation are a challenge, specially if they have to work with non-parallel data, which is a desirable setup in many situations. Voice conversion, in which a model has to impersonate a speaker in a recording, is one of those situations. In this paper, we propose Blow…

2018

Language and Noise Transfer in Speech Enhancement Generative Adversarial Network

ICASSP 2018accepted

Speech enhancement deep learning systems usually require large amounts of training data to operate in broad conditions or real applications. This makes the adaptability of those systems into new, low resource environments an important topic. In this work, we present the results of adapting a speech…

Cited by 0SourceScholar
2017

Effect of acoustic conditions on algorithms to detect Parkinson's disease from speech

ICASSP 2017accepted

Automatic detection of Parkinson's disease (PD) from speech is a basic step towards computer-aided tools supporting the diagnosis and monitoring of the disease. Although several methods have been proposed, their applicability to real-world situations is still unclear. In particular, the effect of ac…

Cited by 0SourceScholar
2016

Discovering rāga motifs by characterizing communities in networks of melodic patterns

ICASSP 2016accepted

Ra̅ga motifs are the main building blocks of the melodic structures in Indian art music. Therefore, the discovery and characterization of such motifs is fundamental for the computational analysis of this music. We propose an approach for discovering ra̅ga motifs from audio music collections. First,…

Cited by 0SourceScholar
2016

Phrase-based rĀga recognition using vector space modeling

ICASSP 2016accepted

Automatic raga recognition is one of the fundamental computational tasks in Indian art music. Motivated by the way seasoned listeners identify ragas, we propose a raga recognition approach based on melodic phrases. Firstly, we extract melodic patterns from a collection of audio recordings in an unsu…

Cited by 0SourceScholar
2015

An evaluation of methodologies for melodic similarity in audio recordings of Indian art music

ICASSP 2015accepted

We perform a comparative evaluation of methodologies for computing similarity between short-time melodic fragments of audio recordings of Indian art music. We experiment with 560 different combinations of procedures and parameter values. These include the choices made for the sampling rate of the me…

Cited by 0SourceScholar