← Search

Tom Bagby

8 accepted papers

2025

Massive Sound Embedding Benchmark (MSEB)

NeurIPS 2025poster

Audio is a critical component of multimodal perception, and any truly intelligent system must demonstrate a wide range of auditory capabilities. These capabilities include transcription, classification, retrieval, reasoning, segmentation, clustering, reranking, and reconstruction. Fundamentally, eac…

Cited by 0SourcecodeScholar
2020

Location-Relative Attention Mechanisms for Robust Long-Form Speech Synthesis

ICASSP 2020accepted

Despite the ability to produce human-level speech for in-domain text, attention-based end-to-end text-to-speech (TTS) systems suffer from text alignment failures that increase in frequency for out-of-domain text. We show that these failures can be addressed using simple location-relative attention m…

Cited by 0SourceScholar
2020

Semi-Supervised Generative Modeling for Controllable Speech Synthesis

ICLR 2020poster

We present a novel generative model that combines state-of-the-art neural text- to-speech (TTS) with semi-supervised probabilistic latent variable models. By providing partial supervision to some of the latent variables, we are able to force them to take on consistent and interpretable purposes, whi…

Cited by 61SourcecodeScholar
2019

Streaming End-to-end Speech Recognition for Mobile Devices

ICASSP 2019accepted

End-to-end (E2E) models, which directly predict output character sequences given input speech, are good candidates for on-device speech recognition. E2E models, however, present numerous challenges: In order to be truly useful, such models must decode speech utterances in a streaming fashion, in rea…

Cited by 677SourceScholar
2018

Sampled Connectionist Temporal Classification

ICASSP 2018accepted

This article introduces and evaluates Sampled Connectionist Temporal Classification (CTC) which connects the CTC criterion to the Cross Entropy (CE) objective through sampling. Instead of computing the logarithm of the sum of the alignment path likelihoods, at each training step the sampled CTC only…

Cited by 0SourceScholar