← Search

Nicholas J. Bryan

12 accepted papers

2026

RETHINKING MUSIC CAPTIONING WITH MUSIC METADATA LLMS

ICASSP 2026poster

Music captioning, or the task of generating a natural language description of music, is useful for both music understanding and controllable music generation. Training captioning models, however, typically requires high-quality music caption data which is scarce compared to metadata (e.g., genre, mo…

Cited by 0SourcePDFScholar
2026

STEMPHONIC: ALL-AT-ONCE FLEXIBLE MULTI-STEM MUSIC GENERATION

ICASSP 2026poster

Music stem generation, the task of producing musically-synchronized and isolated instrument audio clips, offers the potential of greater user control and better alignment with musician workflows compared to conventional text-to-music models. Existing stem generation approaches, however, either rely…

Cited by 0SourcePDFScholar
2025

Presto! Distilling Steps and Layers for Accelerating Music Generation

ICLR 2025spotlight

Despite advances in diffusion-based text-to-music (TTM) methods, efficient, high-quality generation remains a challenge. We introduce Presto!, an approach to inference acceleration for score-based diffusion transformers via reducing both sampling steps and cost per step. To reduce steps, we develop…

Cited by 4SourcePDFScholar
2024

DITTO: Diffusion Inference-Time T-Optimization for Music Generation

ICML 2024oral

We propose Diffusion Inference-Time T-Optimization (DITTO), a general-purpose framework for controlling pre-trained text-to-music diffusion models at inference-time via optimizing initial noise latents. Our method can be used to optimize through any differentiable feature matching loss to achieve a…

2022

Don't Separate, Learn To Remix: End-To-End Neural Remixing With Joint Optimization

ICASSP 2022accepted

The task of manipulating the level and/or effects of individual instruments to recompose a mixture of recordings, or remixing, is common across a variety of applications such as music production, audio-visual post-production, podcasts, and more. This process, however, traditionally requires access t…

Cited by 0SourceScholar
2021

Context-Aware Prosody Correction for Text-Based Speech Editing

ICASSP 2021accepted

Text-based speech editors expedite the process of editing speech recordings by permitting editing via intuitive cut, copy, and paste operations on a speech transcript. A major drawback of current systems, however, is that edited recordings often sound unnatural because of prosody mismatches around e…

Cited by 0SourceScholar
2021

Differentiable Signal Processing With Black-Box Audio Effects

ICASSP 2021accepted

We present a data-driven approach to automate audio signal processing by incorporating stateful third-party, audio effects as layers within a deep neural network. We then train a deep encoder to analyze input audio and control effect parameters to perform the desired signal manipulation, requiring o…

Cited by 0SourceScholar
2021

Few-Shot Continual Learning for Audio Classification

ICASSP 2021accepted

Supervised learning for audio classification typically imposes a fixed class vocabulary, which can be limiting for real-world applications where the target class vocabulary is not known a priori or changes dynamically. In this work, we introduce a few-shot continual learning framework for audio clas…

Cited by 0SourceScholar
2020

Disentangled Multidimensional Metric Learning for Music Similarity

ICASSP 2020accepted

Music similarity search is useful for a variety of creative tasks such as replacing one music recording with another recording with a similar "feel", a common task in video editing. For this task, it is typically necessary to define a similarity metric to compare one recording to another. Music simi…

Cited by 0SourceScholar
2020

Impulse Response Data Augmentation and Deep Neural Networks for Blind Room Acoustic Parameter Estimation

ICASSP 2020accepted

The reverberation time (T60) and the direct-to-reverberant ratio (DRR) are commonly used to characterize room acoustic environments. Both parameters can be measured from an acoustic impulse response (AIR) or using blind estimation methods that perform estimation directly from speech. When neural net…

Cited by 0SourceScholar
2020

One-Shot Parametric Audio Production Style Transfer with Application to Frequency Equalization

ICASSP 2020accepted

Audio production is a difficult process for many people], [and properly manipulating sound to achieve a certain effect is non-trivial. In this paper], [we present a method that facilitates this process by inferring appropriate audio effect parameters in order to make an input recording sound similar…

Cited by 0SourceScholar