← Search

Justin Salamon

24 accepted papers

2026

AUDIOCARDS: STRUCTURED METADATA IMPROVES AUDIO LANGUAGE MODELS FOR SOUND DESIGN

ICASSP 2026oral

Sound designers search for sounds in large sound effects libraries using aspects such as sound class or visual context. However, the metadata needed for such search is often missing or incomplete, and requires significant manual effort to add. Existing solutions to automate this task by generating m…

Cited by 0SourcePDFScholar
2026

AudioChat: Unified Audio Storytelling, Editing, and Understanding with Transfusion Forcing

ICML 2026poster

Despite recent breakthroughs, audio foundation models struggle in processing complex multi-source acoustic scenes. We refer to this challenging domain as audio stories, which can have multiple speakers and background/foreground sound effects. Compared to traditional audio processing tasks, audio sto…

Cited by 0SourcecodeScholar
2025

FLAM: Frame-Wise Language-Audio Modeling

ICML 2025poster

Recent multi-modal audio-language models (ALMs) excel at text-audio retrieval but struggle with frame-wise audio understanding. Prior works use temporal-aware labels or unsupervised training to improve frame-wise capabilities, but they still lack fine-grained labeling capability to pinpoint when an…

Cited by 0SourcePDFScholar
2025

Sketch2Sound: Controllable Audio Generation via Time-Varying Signals and Sonic Imitations

ICASSP 2025accepted

We present Sketch2Sound, a generative audio model capable of creating high-quality sounds from a set of interpretable time-varying control signals: loudness, brightness, and pitch, as well as text prompts. Sketch2Sound can synthesize arbitrary sounds from sonic imitations (i.e., a vocal imitation or…

Cited by 0SourceScholar
2025

Video-Guided Foley Sound Generation with Multimodal Controls

CVPR 2025poster

Generating sound effects for videos often requires creating artistic sound effects that diverge significantly from real-life sources and flexible control in the sound design. To address this problem, we introduce *MultiFoley*, a model designed for video-guided sound generation that supports multimod…

Cited by 10SourcePDFScholar
2023

Conditional Generation of Audio From Video via Foley Analogies

CVPR 2023poster

The sound effects that designers add to videos are designed to convey a particular artistic effect and, thus, may be quite different from a scene's true sound. Inspired by the challenges of creating a soundtrack for a video that differs from its true sound, but that nonetheless matches the actions o…

2023

Language-Guided Audio-Visual Source Separation via Trimodal Consistency

CVPR 2023poster

We propose a self-supervised approach for learning to perform audio source separation in videos based on natural language queries, using only unlabeled video and audio pairs as training data. A key challenge in this task is learning to associate the linguistic description of a sound-emitting object…

Cited by 19SourcePDFScholar
2023

Language-Guided Music Recommendation for Video via Prompt Analogies

CVPR 2023highlight

We propose a method to recommend music for an input video while allowing a user to guide music selection with free-form natural language. A key challenge of this problem setting is that existing music video datasets provide the needed (video, music) training pairs, but lack text descriptions of the…

Cited by 30SourcePDFScholar
2021

Few-Shot Continual Learning for Audio Classification

ICASSP 2021accepted

Supervised learning for audio classification typically imposes a fixed class vocabulary, which can be limiting for real-world applications where the target class vocabulary is not known a priori or changes dynamically. In this work, we introduce a few-shot continual learning framework for audio clas…

Cited by 0SourceScholar
2021

Sound Event Detection and Separation: A Benchmark on Desed Synthetic Soundscapes

ICASSP 2021accepted

We propose a benchmark of state-of-the-art sound event detection systems (SED). We design synthetic evaluation sets to focus on specific sound event detection challenges. We analyze the performance of the submissions to DCASE 2020 Task 4 as a function of time-related modifications (time position of…

Cited by 0SourceScholar
2021

What's all the Fuss about Free Universal Sound Separation Data?

ICASSP 2021accepted

We introduce the Free Universal Sound Separation (FUSS) dataset, a new corpus for experiments in separating mixtures of an unknown number of sounds from an open domain of sound types. The dataset consists of 23 hours of single-source audio data drawn from 357 classes, which are used to create mixtur…

Cited by 0SourceScholar
2020

Chirping up the Right Tree: Incorporating Biological Taxonomies into Deep Bioacoustic Classifiers

ICASSP 2020accepted

Class imbalance in the training data hinders the generalization ability of machine listening systems. In the context of bioacoustics, this issue may be circumvented by aggregating species labels into super-groups of higher taxonomic rank: genus, family, order, and so forth. However, different applic…

Cited by 0SourceScholar
2020

Disentangled Multidimensional Metric Learning for Music Similarity

ICASSP 2020accepted

Music similarity search is useful for a variety of creative tasks such as replacing one music recording with another recording with a similar "feel", a common task in video editing. For this task, it is typically necessary to define a similarity metric to compare one recording to another. Music simi…

Cited by 0SourceScholar
2020

Sound Event Detection in Synthetic Domestic Environments

ICASSP 2020accepted

We present a comparative analysis of the performance of state-of-the-art sound event detection systems. In particular, we study the robustness of the systems to noise and signal degradation, which is known to impact model generalization. Our analysis is based on the results of task 4 of the DCASE 20…

Cited by 0SourceScholar
2019

Look, Listen, and Learn More: Design Choices for Deep Audio Embeddings

ICASSP 2019accepted

A considerable challenge in applying deep learning to audio classification is the scarcity of labeled data. An increasingly popular solution is to learn deep audio embeddings from large audio collections and use them to train shallow classifiers using small labeled datasets. Look, Listen, and Learn…

Cited by 0SourceScholar
2018

Birdvox-Full-Night: A Dataset and Benchmark for Avian Flight Call Detection

ICASSP 2018accepted

This article addresses the automatic detection of vocal, nocturnally migrating birds from a network of acoustic sensors. Thus far, owing to the lack of annotated continuous recordings, existing methods had been benchmarked in a binary classification setting (presence vs. absence). Instead, with the…

Cited by 0SourceScholar
2018

Crepe: A Convolutional Representation for Pitch Estimation

ICASSP 2018accepted

The task of estimating the fundamental frequency of a monophonic sound recording, also known as pitch tracking, is fundamental to audio processing with multiple applications in speech processing and music information retrieval. To date, the best performing techniques, such as the pYIN algorithm, are…

Cited by 0SourceScholar
2018

Investigating the Effect of Sound-Event Loudness on Crowdsourced Audio Annotations

ICASSP 2018accepted

Audio annotation is an important step in developing machine-listening systems. It is also a time consuming process, which has motivated investigators to crowdsource audio annotations. However, there are many factors that affect annotations, many of which have not been adequately investigated. In pre…

Cited by 0SourceScholar
2017

Fusing shallow and deep learning for bioacoustic bird species classification

ICASSP 2017accepted

Automated classification of organisms to species based on their vocalizations would contribute tremendously to abilities to monitor biodiversity, with a wide range of applications in the field of ecology. In particular, automated classification of migrating birds' flight calls could yield new biolog…

Cited by 0SourceScholar