← Search

Xavier Serra

24 accepted papers

2026

Benchmarking Music Autotagging with MGPHot Expert Annotations vs. Generic Tag Datasets

ICASSP 2026oral

Music autotagging aims to automatically assign descriptive tags, such as genre, mood, or instrumentation, to audio recordings. Due to its challenges, diversity of semantic descriptions, and practical value in various applications, it has become a common downstream task for evaluating the performance…

Cited by 0SourcePDFScholar
2023

Flowgrad: Using Motion for Visual Sound Source Localization

ICASSP 2023accepted

Most recent work in visual sound source localization relies on semantic audio-visual representations learned in a self-supervised manner and, by design, excludes temporal information present in videos. While it proves to be effective for widely used benchmark datasets, the method falls short for cha…

Cited by 0SourceScholar
2023

Pre-Training Strategies Using Contrastive Learning and Playlist Information for Music Classification and Similarity

ICASSP 2023accepted

In this work, we investigate an approach that relies on contrastive learning and music metadata as a weak source of supervision to train music representation models. Recent studies show that contrastive learning can be used with editorial metadata (e.g., artist or album name) to learn audio represen…

Cited by 0SourceScholar
2022

Score Difficulty Analysis for Piano Performance Education based on Fingering

ICASSP 2022accepted

In this paper, we introduce score difficulty classification as a sub-task of music information retrieval (MIR), which may be used in music education technologies, for personalised curriculum generation, and score retrieval. We introduce a novel dataset for our task, Mikrokosmos-difficulty, containin…

Cited by 0SourceScholar
2022

Urban Sound & Sight: Dataset And Benchmark For Audio-Visual Urban Scene Understanding

ICASSP 2022accepted

Automatic audio-visual urban traffic understanding is a growing area of research with many potential applications of value to industry, academia, and the public sector. Yet, the lack of well-curated resources for training and evaluating models to research in this area hinders their development. To a…

Cited by 16SourceScholar
2021

Learning Contextual Tag Embeddings for Cross-Modal Alignment of Audio and Tags

ICASSP 2021accepted

Self-supervised audio representation learning offers an attractive alternative for obtaining generic audio embeddings, capable to be employed into various downstream tasks. Published approaches that consider both audio and words/tags associated with audio do not employ text processing models that ar…

Cited by 0SourceScholar
2021

Loopnet: Musical Loop Synthesis Conditioned on Intuitive Musical Parameters

ICASSP 2021accepted

Loops, seamlessly repeatable musical segments, are a cornerstone of modern music production. Contemporary artists often mix and match various sampled or pre-recorded loops based on musical criteria such as rhythm, harmony and timbral texture to create compositions. Taking such criteria into account,…

Cited by 0SourceScholar
2021

Melon Playlist Dataset: A Public Dataset for Audio-Based Playlist Generation and Music Tagging

ICASSP 2021accepted

One of the main limitations in the field of audio signal processing is the lack of large public datasets with audio representations and high-quality annotations due to restrictions of copyrighted commercial music. We present Melon Playlist Dataset, a public dataset of mel-spectrograms for 649,091 tr…

Cited by 0SourceScholar
2021

Multimodal Metric Learning for Tag-Based Music Retrieval

ICASSP 2021accepted

Tag-based music retrieval is crucial to browse large-scale mu-sic libraries efficiently. Hence, automatic music tagging has been actively explored, mostly as a classification task, which has an inherent limitation: a fixed vocabulary. On the other hand, metric learning enables flexible vocabularies…

Cited by 0SourceScholar
2021

Unsupervised Contrastive Learning of Sound Event Representations

ICASSP 2021accepted

Self-supervised representation learning can mitigate the limitations in recognition tasks with few manually labeled data but abundant unlabeled data—a common scenario in sound event research. In this work, we explore unsupervised contrastive learning as a way to learn sound event representations. To…

Cited by 0SourceScholar
2020

Neural Percussive Synthesis Parameterised by High-Level Timbral Features

ICASSP 2020accepted

We present a deep neural network-based methodology for synthesising percussive sounds with control over high-level timbral characteristics of the sounds. This approach allows for intuitive control of a synthesizer, enabling the user to shape sounds without extensive knowledge of signal processing. W…

Cited by 26SourceScholar
2019

Learning Sound Event Classifiers from Web Audio with Noisy Labels

ICASSP 2019accepted

As sound event classification moves towards larger datasets, issues of label noise become inevitable. Web sites can supply large volumes of user-contributed audio and metadata, but inferring labels from this metadata introduces errors due to unreliable inputs, and limitations in the mapping. There i…

Cited by 0SourceScholar
2017

Designing efficient architectures for modeling temporal features with convolutional neural networks

ICASSP 2017accepted

Many researchers use convolutional neural networks with small rectangular filters for music (spectrograms) classification. First, we discuss why there is no reason to use this filters setup by default and second, we point that more efficient architectures could be implemented if the characteristics…

Cited by 0SourceScholar
2016

A generalized Bayesian model for tracking long metrical cycles in acoustic music signals

ICASSP 2016accepted

Most musical phenomena involve repetitive structures that enable listeners to track meter, i.e. the tactus or beat, the longer over-arching measure or bar, and possibly other related layers. Meters with long measure duration, sometimes lasting more than a minute, occur in many music cultures, e.g. f…

Cited by 0SourceScholar
2016

Discovering rāga motifs by characterizing communities in networks of melodic patterns

ICASSP 2016accepted

Ra̅ga motifs are the main building blocks of the melodic structures in Indian art music. Therefore, the discovery and characterization of such motifs is fundamental for the computational analysis of this music. We propose an approach for discovering ra̅ga motifs from audio music collections. First,…

Cited by 0SourceScholar
2016

Phrase-based rĀga recognition using vector space modeling

ICASSP 2016accepted

Automatic raga recognition is one of the fundamental computational tasks in Indian art music. Motivated by the way seasoned listeners identify ragas, we propose a raga recognition approach based on melodic phrases. Firstly, we extract melodic patterns from a collection of audio recordings in an unsu…

Cited by 0SourceScholar
2015

An evaluation of methodologies for melodic similarity in audio recordings of Indian art music

ICASSP 2015accepted

We perform a comparative evaluation of methodologies for computing similarity between short-time melodic fragments of audio recordings of Indian art music. We experiment with 560 different combinations of procedures and parameter values. These include the choices made for the sampling rate of the me…

Cited by 0SourceScholar