← Search

Kyle Kastner

10 accepted papers

2025

Audio Diffusion with Large Language Models

ICASSP 2025accepted

In this paper, we explore an alternate approach to the popular method of using large language models (LLMs) as a second decoder for Automated Speech Recognition (ASR) and speech understanding tasks. We propose to employ diffusion networks to generate a correction signal that can be applied on the or…

Cited by 0SourceScholar
2025

Speech Re-Painting for Robust ASR

ICASSP 2025accepted

Synthetic speech is a useful source for augmentation of automatic speech recognition (ASR) systems, but there is a "sim-to-real" gap between synthetic and real speech that can limit generalization. The natural variability of real speech is essential to the training of robust ASR systems. While synth…

Cited by 0SourceScholar
2024

Adaptive Accompaniment with ReaLchords

ICML 2024poster

Jamming requires coordination, anticipation, and collaborative creativity between musicians. Current generative models of music produce expressive output but are not able to generate in an online manner, meaning simultaneously with other musicians (human or otherwise). We propose ReaLchords, an onli…

Cited by 3SourcePDFScholar
2024

Extending Multilingual Speech Synthesis to 100+ Languages without Transcribed Data

ICASSP 2024accepted

Collecting high-quality studio recordings of audio is challenging, which limits the language coverage of text-to-speech (TTS) systems. This paper proposes a framework for scaling a multilingual TTS model to 100+ languages using found data without supervision. The proposed framework combines speech-t…

Cited by 0SourceScholar
2023

Understanding Shared Speech-Text Representations

ICASSP 2023accepted

Recently, a number of approaches to train speech models by incorporating text into end-to-end models have been developed, with Maestro advancing state-of-the-art automatic speech recognition (ASR) and Speech Translation (ST) performance. In this paper, we expand our understanding of the resulting sh…

Cited by 0SourceScholar
2022

MIDI-DDSP: Detailed Control of Musical Performance via Hierarchical Modeling

ICLR 2022oral

Musical expression requires control of both what notes that are played, and how they are performed. Conventional audio synthesizers provide detailed expressive controls, but at the cost of realism. Black-box neural audio synthesis and concatenative samplers can produce realistic audio, but have few…

2017

Learning to Discover Sparse Graphical Models

ICML 2017poster

We consider structure discovery of undirected graphical models from observational data. Inferring likely structures from few examples is a complex task often requiring the formulation of priors and sophisticated inference procedures. Popular methods rely on estimating a penalized maximum likelihood…

Cited by 40SourcePDFScholar
2015

A Recurrent Latent Variable Model for Sequential Data

NeurIPS 2015poster

In this paper, we explore the inclusion of latent random variables into the hidden state of a recurrent neural network (RNN) by combining the elements of the variational autoencoder. We argue that through the use of high-level latent random variables, the variational RNN (VRNN) can model the kind of…