← Search

Magdalena Fuentes

8 accepted papers

2025

A Critical Assessment of Visual Sound Source Localization Models Including Negative Audio

ICASSP 2025accepted

The task of Visual Sound Source Localization (VSSL) involves identifying the location of sound sources in visual scenes, integrating audio-visual data for enhanced scene understanding. Despite advancements in state-of-the-art (SOTA) models, we observe three critical flaws: i) The evaluation of the m…

Cited by 0SourceScholar
2025

Twenty-Five Years of MIR Research: Achievements, Practices, Evaluations, and Future Challenges

ICASSP 2025accepted

In this paper, we trace the evolution of Music Information Retrieval (MIR) over the past 25 years. While MIR gathers all kinds of research related to music informatics, a large part of it focuses on signal processing techniques for music data, fostering a close relationship with the IEEE Audio and A…

Cited by 1SourceScholar
2023

Does a Quieter City Mean Fewer Complaints? The Sounds of New York City During Covid-19 Lockdown

ICASSP 2023accepted

The COVID-19 pandemic had an unprecedented effect in human activity and city landscapes. A very notorious transformation during this period was the change in noise levels and patterns across cities. Small scale studies have show this change in noise levels across different locations in the globe. In…

Cited by 0SourceScholar
2023

Flowgrad: Using Motion for Visual Sound Source Localization

ICASSP 2023accepted

Most recent work in visual sound source localization relies on semantic audio-visual representations learned in a self-supervised manner and, by design, excludes temporal information present in videos. While it proves to be effective for widely used benchmark datasets, the method falls short for cha…

Cited by 0SourceScholar
2023

Tempo vs. Pitch: Understanding Self-Supervised Tempo Estimation

ICASSP 2023accepted

Self-supervision methods learn representations by solving pretext tasks that do not require human-generated labels, alleviating the need for time-consuming annotations. These methods have been applied in computer vision, natural language processing, environmental sound analysis, and recently in musi…

Cited by 0SourceScholar
2022

Urban Sound & Sight: Dataset And Benchmark For Audio-Visual Urban Scene Understanding

ICASSP 2022accepted

Automatic audio-visual urban traffic understanding is a growing area of research with many potential applications of value to industry, academia, and the public sector. Yet, the lack of well-curated resources for training and evaluating models to research in this area hinders their development. To a…

Cited by 16SourceScholar
2019

A Music Structure Informed Downbeat Tracking System Using Skip-chain Conditional Random Fields and Deep Learning

ICASSP 2019accepted

In recent years the task of downbeat tracking has received increasing attention and the state of the art has been improved with the introduction of deep learning methods. Among proposed solutions, existing systems exploit short-term musical rules as part of their language modelling. In this work we…

Cited by 0SourceScholar