← Search

Juan Pablo Bello

29 accepted papers

2026

Controllable Embedding Transformation for Mood-Guided Music Retrieval

ICASSP 2026poster

Music representations are the backbone of modern recommendation systems, powering playlist generation, similarity search, and personalized discovery. Yet most embeddings offer little control for adjusting a single musical attribute, e.g., changing only the mood of a track while preserving its genre…

Cited by 1SourcePDFScholar
2024

Spatial Scaper: A Library to Simulate and Augment Soundscapes for Sound Event Localization and Detection in Realistic Rooms

ICASSP 2024accepted

Sound event localization and detection (SELD) is an important task in machine listening. Major advancements rely on simulated data with sound events in specific rooms and strong spatio-temporal labels. SELD data is simulated by convolving spatialy-localized room impulse responses (RIRs) with sound w…

Cited by 0SourceScholar
2023

Does a Quieter City Mean Fewer Complaints? The Sounds of New York City During Covid-19 Lockdown

ICASSP 2023accepted

The COVID-19 pandemic had an unprecedented effect in human activity and city landscapes. A very notorious transformation during this period was the change in noise levels and patterns across cities. Small scale studies have show this change in noise levels across different locations in the globe. In…

Cited by 0SourceScholar
2023

Exploring Approaches to Multi-Task Automatic Synthesizer Programming

ICASSP 2023accepted

Automatic Synthesizer Programming is the task of transforming an audio signal that was generated from a virtual instrument, into the parameters of a sound synthesizer that would generate this signal. In the past, this could only be done for one virtual instrument. In this paper, we expand the curren…

Cited by 0SourceScholar
2023

Flowgrad: Using Motion for Visual Sound Source Localization

ICASSP 2023accepted

Most recent work in visual sound source localization relies on semantic audio-visual representations learned in a self-supervised manner and, by design, excludes temporal information present in videos. While it proves to be effective for widely used benchmark datasets, the method falls short for cha…

Cited by 0SourceScholar
2022

Urban Sound & Sight: Dataset And Benchmark For Audio-Visual Urban Scene Understanding

ICASSP 2022accepted

Automatic audio-visual urban traffic understanding is a growing area of research with many potential applications of value to industry, academia, and the public sector. Yet, the lack of well-curated resources for training and evaluating models to research in this area hinders their development. To a…

Cited by 16SourceScholar
2022

Wav2CLIP: Learning Robust Audio Representations from Clip

ICASSP 2022accepted

We propose Wav2CLIP, a robust audio representation learning method by distilling from Contrastive Language-Image Pre-training (CLIP). We systematically evaluate Wav2CLIP on a variety of audio tasks including classification, retrieval, and generation, and show that Wav2CLIP can outperform several pub…

Cited by 0SourceScholar
2021

Few-Shot Continual Learning for Audio Classification

ICASSP 2021accepted

Supervised learning for audio classification typically imposes a fixed class vocabulary, which can be limiting for real-world applications where the target class vocabulary is not known a priori or changes dynamically. In this work, we introduce a few-shot continual learning framework for audio clas…

Cited by 0SourceScholar
2021

Multi-Task Self-Supervised Pre-Training for Music Classification

ICASSP 2021accepted

Deep learning is very data hungry, and supervised learning especially requires massive labeled data to work well. Machine listening research often suffers from limited labeled data problem, as human annotations are costly to acquire, and annotations for audio are time consuming and less intuitive. B…

Cited by 0SourceScholar
2021

Specialized Embedding Approximation for Edge Intelligence: A Case Study in Urban Sound Classification

ICASSP 2021accepted

Embedding models that encode semantic information into low-dimensional vector representations are useful in various machine learning tasks with limited training data. However, these models are typically too large to support inference in small edge devices, which motivates training of smaller yet com…

Cited by 0SourceScholar
2020

Chirping up the Right Tree: Incorporating Biological Taxonomies into Deep Bioacoustic Classifiers

ICASSP 2020accepted

Class imbalance in the training data hinders the generalization ability of machine listening systems. In the context of bioacoustics, this issue may be circumvented by aggregating species labels into super-groups of higher taxonomic rank: genus, family, order, and so forth. However, different applic…

Cited by 0SourceScholar
2020

Learning the Helix Topology of Musical Pitch

ICASSP 2020accepted

To explain the consonance of octaves, music psychologists represent pitch as a helix where azimuth and axial coordinate correspond to pitch class and pitch height respectively. This article addresses the problem of discovering this helical structure from unlabeled audio data. We measure Pearson corr…

Cited by 0SourceScholar
2019

A Music Structure Informed Downbeat Tracking System Using Skip-chain Conditional Random Fields and Deep Learning

ICASSP 2019accepted

In recent years the task of downbeat tracking has received increasing attention and the state of the art has been improved with the introduction of deep learning methods. Among proposed solutions, existing systems exploit short-term musical rules as part of their language modelling. In this work we…

Cited by 0SourceScholar
2019

Active Learning for Efficient Audio Annotation and Classification with a Large Amount of Unlabeled Data

ICASSP 2019accepted

There are many sound classification problems that have target classes which are rare or unique to the context of the problem. For these problems, existing data sets are not sufficient and we must create new problem-specific datasets to train classification models. However, annotating a new dataset f…

Cited by 0SourceScholar
2019

Look, Listen, and Learn More: Design Choices for Deep Audio Embeddings

ICASSP 2019accepted

A considerable challenge in applying deep learning to audio classification is the scarcity of labeled data. An increasingly popular solution is to learn deep audio embeddings from large audio collections and use them to train shallow classifiers using small labeled datasets. Look, Listen, and Learn…

Cited by 0SourceScholar
2018

Birdvox-Full-Night: A Dataset and Benchmark for Avian Flight Call Detection

ICASSP 2018accepted

This article addresses the automatic detection of vocal, nocturnally migrating birds from a network of acoustic sensors. Thus far, owing to the lack of annotated continuous recordings, existing methods had been benchmarked in a binary classification setting (presence vs. absence). Instead, with the…

Cited by 0SourceScholar
2018

Crepe: A Convolutional Representation for Pitch Estimation

ICASSP 2018accepted

The task of estimating the fundamental frequency of a monophonic sound recording, also known as pitch tracking, is fundamental to audio processing with multiple applications in speech processing and music information retrieval. To date, the best performing techniques, such as the pYIN algorithm, are…

Cited by 0SourceScholar
2018

Investigating the Effect of Sound-Event Loudness on Crowdsourced Audio Annotations

ICASSP 2018accepted

Audio annotation is an important step in developing machine-listening systems. It is also a time consuming process, which has motivated investigators to crowdsource audio annotations. However, there are many factors that affect annotations, many of which have not been adequately investigated. In pre…

Cited by 0SourceScholar
2017

Fusing shallow and deep learning for bioacoustic bird species classification

ICASSP 2017accepted

Automated classification of organisms to species based on their vocalizations would contribute tremendously to abilities to monitor biodiversity, with a wide range of applications in the field of ecology. In particular, automated classification of migrating birds' flight calls could yield new biolog…

Cited by 0SourceScholar
2017

Towards the characterization of singing styles in world music

ICASSP 2017accepted

In this paper we focus on the characterization of singing styles in world music.We develop a set of contour features capturing pitch structure and melodic embellishments.Using these features we train a binary classifier to distinguish vocal from non-vocal contours and learn a dictionary of singing s…

Cited by 0SourceScholar
2016

Feature adapted convolutional neural networks for downbeat tracking

ICASSP 2016accepted

We define a novel system for the automatic estimation of downbeat positions from audio music signals. New rhythm and melodic features are introduced and feature adapted convolutional neural networks are used to take advantage of their specificity. Indeed, invariance to melody transposition, chroma d…

Cited by 23SourceScholar
2015

Downbeat tracking with multiple features and deep neural networks

ICASSP 2015accepted

In this paper, we introduce a novel method for the automatic estimation of downbeat positions from music signals. Our system relies on the computation of musically inspired features capturing important aspects of music such as timbre, harmony, rhythmic patterns, or local similarities in both timbre…

Cited by 39SourceScholar