← Search

RJ Skerry-Ryan

5 accepted papers

2025

Long-Form Speech Generation with Spoken Language Models

ICML 2025oral

We consider the generative modeling of speech over multiple minutes, a requirement for long-form multimedia generation and audio-native voice assistants. However, textless spoken language models struggle to generate plausible speech past tens of seconds, due to high temporal resolution of speech tok…

2025

Robust and Unbounded Length Generalization in Autoregressive Transformer-Based Text-to-Speech

NAACL 2025long

Autoregressive (AR) Transformer-based sequence models are known to have difficulty generalizing to sequences longer than those seen during training. When applied to text-to-speech (TTS), these models tend to drop or repeat words or produce erratic output, especially for longer utterances. In this pa…

2024

Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

ICLR 2024poster

We present Spectron, a novel approach to adapting pre-trained large language models (LLMs) to perform spoken question answering (QA) and speech continuation. By endowing the LLM with a pre-trained speech encoder, our model becomes able to take speech inputs and generate speech outputs. The entire sy…

Cited by 40SourcePDFScholar
2020

Semi-Supervised Generative Modeling for Controllable Speech Synthesis

ICLR 2020poster

We present a novel generative model that combines state-of-the-art neural text- to-speech (TTS) with semi-supervised probabilistic latent variable models. By providing partial supervision to some of the latent variables, we are able to force them to take on consistent and interpretable purposes, whi…

Cited by 61SourcecodeScholar
2018

Towards End-to-End Prosody Transfer for Expressive Speech Synthesis with Tacotron

ICML 2018oral

We present an extension to the Tacotron speech synthesis architecture that learns a latent embedding space of prosody, derived from a reference acoustic representation containing the desired prosody. We show that conditioning Tacotron on this learned embedding space results in synthesized audio that…

Cited by 749SourcePDFScholar