← Search

Bryan Pardo

20 accepted papers

2025

Code Drift: Towards Idempotent Neural Audio Codecs

ICASSP 2025accepted

Neural codecs have demonstrated strong performance in high-fidelity compression of audio signals at low bitrates. The token-based representations produced by these codecs have proven particularly useful for generative modeling. While much research has focused on improvements in compression ratio and…

Cited by 0SourceScholar
2025

Sketch2Sound: Controllable Audio Generation via Time-Varying Signals and Sonic Imitations

ICASSP 2025accepted

We present Sketch2Sound, a generative audio model capable of creating high-quality sounds from a set of interpretable time-varying control signals: loudness, brightness, and pitch, as well as text prompts. Sketch2Sound can synthesize arbitrary sounds from sonic imitations (i.e., a vocal imitation or…

Cited by 0SourceScholar
2025

Text2FX: Harnessing CLAP Embeddings for Text-Guided Audio Effects

ICASSP 2025accepted

This work introduces Text2FX, a method that leverages CLAP embeddings and differentiable digital signal processing to control audio effects, such as equalization and reverberation, using open-vocabulary natural language prompts (e.g., "make this sound in-your-face and bold"). Text2FX operates withou…

Cited by 0SourceScholar
2025

Towards A Translative Model of Sperm Whale Vocalization

NeurIPS 2025poster

Sperm whales communicate in short sequences of clicks known as codas. We present WhAM (Whale Acoustics Model), the first transformer-based model capable of generating synthetic sperm whale codas from any audio prompt. WhAM is built by finetuning VampNet, a masked acoustic token model pretrained on m…

Cited by 1SourcecodeScholar
2024

Crowdsourced and Automatic Speech Prominence Estimation

ICASSP 2024accepted

The prominence of a spoken word is the degree to which an average native listener perceives the word as salient or emphasized relative to its context. Speech prominence estimation is the process of assigning a numeric value to the prominence of each word in an utterance. These prominence labels are…

Cited by 0SourceScholar
2022

Effective and Inconspicuous Over-the-Air Adversarial Examples with Adaptive Filtering

ICASSP 2022accepted

While deep neural networks achieve state-of-the-art performance on many audio classification tasks, they are known to be vulnerable to adversarial examples - artificially-generated perturbations of natural instances that cause a network to make incorrect predictions. In this work we demonstrate a no…

Cited by 0SourceScholar
2022

Improving Source Separation by Explicitly Modeling Dependencies between Sources

ICASSP 2022accepted

We propose a new method for training a supervised source separation system that aims to learn the interdependent relationships between all combinations of sources in a mixture. Rather than independently estimating each source from a mix, we reframe the source separation problem as an Orderless Neura…

Cited by 0SourceScholar
2022

VoiceBlock: Privacy through Real-Time Adversarial Attacks with Audio-to-Audio Models

NeurIPS 2022accept

As governments and corporations adopt deep learning systems to collect and analyze user-generated audio data, concerns about security and privacy naturally emerge in areas such as automatic speaker recognition. While audio adversarial examples offer one route to mislead or evade these invasive syste…

2021

Context-Aware Prosody Correction for Text-Based Speech Editing

ICASSP 2021accepted

Text-based speech editors expedite the process of editing speech recordings by permitting editing via intuitive cut, copy, and paste operations on a speech transcript. A major drawback of current systems, however, is that edited recordings often sound unnatural because of prosody mismatches around e…

Cited by 0SourceScholar
2020

Simultaneous Separation and Transcription of Mixtures with Multiple Polyphonic and Percussive Instruments

ICASSP 2020accepted

We present a single deep learning architecture that can both separate an audio recording of a musical mixture into constituent single-instrument recordings and transcribe these instruments into a human-readable format at the same time, learning a shared musical representation for both tasks. This no…

Cited by 0SourceScholar
2019

Bootstrapping Single-channel Source Separation via Unsupervised Spatial Clustering on Stereo Mixtures

ICASSP 2019accepted

Separating an audio scene into isolated sources is a fundamental problem in computer audition, analogous to image segmentation in visual scene analysis. Source separation systems based on deep learning are currently the most successful approaches for solving the underdetermined separation problem, w…

Cited by 0SourceScholar
2018

Blind Estimation of the Speech Transmission Index for Speech Quality Prediction

ICASSP 2018accepted

The speech transmission index (STI) of a listening position within a given room indicates the quality and intelligibility of speech uttered in that room. The measure is very reliable for predicting speech intelligibility in many room conditions but requires an STI measurement of the impulse response…

Cited by 0SourceScholar
2016

Fast and easy crowdsourced perceptual audio evaluation

ICASSP 2016accepted

Automated objective methods of audio evaluation are fast, cheap, and require little effort by the investigator. However, objective evaluation methods do not exist for the output of all audio processing algorithms, often have output that correlates poorly with human quality assessments, and require g…

Cited by 0SourceScholar
2015

A simple user interface system for recovering patterns repeating in time and frequency in mixtures of sounds

ICASSP 2015accepted

Repetition is a fundamental element in generating and perceiving structure in audio. Especially in music, structures tend to be composed of patterns that repeat through time (e.g., rhythmic elements in a musical accompaniment), and also frequency (e.g., different notes of the same instrument). The a…

Cited by 0SourceScholar