← Search

Jonah Casebeer

11 accepted papers

2025

Presto! Distilling Steps and Layers for Accelerating Music Generation

ICLR 2025spotlight

Despite advances in diffusion-based text-to-music (TTM) methods, efficient, high-quality generation remains a challenge. We introduce Presto!, an approach to inference acceleration for score-based diffusion transformers via reducing both sampling steps and cost per step. To reduce steps, we develop…

Cited by 4SourcePDFScholar
2025

REGEN: Learning Compact Video Embedding with (Re-)Generative Decoder

ICCV 2025poster

We present a novel perspective on learning video embedders for generative modeling: rather than requiring an exact reproduction of an input video, an effective embedder should focus on synthesizing visually plausible reconstructions. This relaxed criterion enables substantial improvements in compres…

Cited by 0SourcePDFScholar
2021

Communication-Cost Aware Microphone Selection for Neural Speech Enhancement with Ad-Hoc Microphone Arrays

ICASSP 2021accepted

In this paper, we present a method for jointly-learning a microphone selection mechanism and a speech enhancement network for multi-channel speech enhancement with an ad-hoc microphone array. The attention-based microphone selection mechanism is trained to reduce communication-costs through a penalt…

Cited by 0SourceScholar
2021

Enhancing into the Codec: Noise Robust Speech Coding with Vector-Quantized Autoencoders

ICASSP 2021accepted

Audio codecs based on discretized neural autoencoders have recently been developed and shown to provide significantly higher compression levels for comparable quality speech out-put. However, these models are tightly coupled with speech content, and produce unintended outputs in noisy conditions. Ba…

Cited by 0SourceScholar
2020

Efficient Trainable Front-Ends for Neural Speech Enhancement

ICASSP 2020accepted

Many neural speech enhancement and source separation systems operate in the time-frequency domain. Such models often benefit from making their Short-Time Fourier Transform (STFT) front-ends trainable. In current literature, these are implemented as large Discrete Fourier Transform matrices; which ar…

Cited by 4SourceScholar
2019

Dimensional Analysis of Laughter in Female Conversational Speech

ICASSP 2019accepted

How do people hear laughter in expressive, unprompted speech? What is the range of expressivity and function of laughter in this speech, and how can laughter inform the recognition of higher-level expressive dimensions in a corpus? This paper presents a scalable method for collecting natural human d…

Cited by 0SourceScholar
2018

Verbal Protest Recognition in Children with Autism

ICASSP 2018accepted

Real-time detection of verbal protest (sensory overload-induced crying) in children with autism is a first step towards understanding the precursors of challenging behaviors associated with autism. Detection of verbal protest is useful for both autism researchers interested in exploring just-in-time…

Cited by 0SourceScholar