← Search

Jordi Pons

21 accepted papers

2026

LOW-RESOURCE GUIDANCE FOR CONTROLLABLE LATENT AUDIO DIFFUSION

ICASSP 2026poster

Generative audio requires fine-grained controllable outputs, yet most existing methods require model retraining on specific controls or inference-time controls (\textit{e.g.}, guidance) that can also be computationally demanding. By examining the bottlenecks of existing guidance-based controls, in p…

Cited by 0SourcePDFScholar
2025

Scaling Transformers for Low-Bitrate High-Quality Speech Coding

ICLR 2025poster

The tokenization of audio with neural audio codec models is a vital part of modern AI pipelines for the generation or understanding of speech, alone or in a multimodal context. Traditionally such tokenization models have concentrated on low parameter-count architectures using only components with st…

2024

GASS: Generalizing Audio Source Separation with Large-Scale Data

ICASSP 2024accepted

Universal source separation targets at separating the audio sources of an arbitrary mix, removing the constraint to operate on a specific domain like speech or music. Yet, the potential of universal source separation is limited because most existing works focus on mixes with predominantly sound even…

Cited by 0SourceScholar
2023

Adversarial Permutation Invariant Training for Universal Sound Separation

ICASSP 2023accepted

Universal sound separation consists of separating mixes with arbitrary sounds of different types, and permutation invariant training (PIT) is used to train source agnostic models that do so. In this work, we complement PIT with adversarial losses but find it challenging with the standard formulation…

Cited by 0SourceScholar
2023

Full-Band General Audio Synthesis with Score-Based Diffusion

ICASSP 2023accepted

Recent works have shown the capability of deep generative models to tackle general audio synthesis from a single label, producing a variety of impulsive, tonal, and environmental sounds. Such models operate on band-limited signals and, as a result of an autoregressive approach, they are typically co…

Cited by 0SourceScholar
2022

On Loss Functions and Evaluation Metrics for Music Source Separation

ICASSP 2022accepted

We investigate which loss functions provide better separations via benchmarking an extensive set of those for music source separation. To that end, we first survey the most representative audio source separation losses we identified, to later consistently benchmark them in a controlled experimental…

Cited by 0SourceScholar
2022

Pixinwav: Residual Steganography for Hiding Pixels in Audio

ICASSP 2022accepted

Steganography comprises the mechanics of hiding data in a host media that may be publicly available. While previous works focused on unimodal setups (e.g., hiding images in images, or hiding audio in audio), PixInWav targets the multimodal case of hiding images in audio. To this end, we propose a no…

Cited by 0SourceScholar
2021

Automatic Multitrack Mixing With A Differentiable Mixing Console Of Neural Audio Effects

ICASSP 2021accepted

Applications of deep learning to automatic multitrack mixing are largely unexplored. This is partly due to the limited available data, coupled with the fact that such data is relatively unstructured and variable. To address these challenges, we propose a domain-inspired model with a strong inductive…

Cited by 0SourceScholar
2017

Designing efficient architectures for modeling temporal features with convolutional neural networks

ICASSP 2017accepted

Many researchers use convolutional neural networks with small rectangular filters for music (spectrograms) classification. First, we discuss why there is no reason to use this filters setup by default and second, we point that more efficient architectures could be implemented if the characteristics…

Cited by 0SourceScholar
2015

On automatic drum transcription using non-negative matrix deconvolution and itakura saito divergence

ICASSP 2015accepted

This paper presents an investigation into the detection and classification of drum sounds in polyphonic music and drum loops using non-negative matrix deconvolution (NMD) and the Itakura Saito divergence. The Itakura Saito divergence has recently been proposed as especially appropriate for decomposi…

Cited by 0SourceScholar