← Search

Joel Shor

5 accepted papers

2022

Universal Paralinguistic Speech Representations Using self-Supervised Conformers

ICASSP 2022accepted

Many speech applications require understanding aspects beyond the words being spoken, such as recognizing emotion, detecting whether the speaker is wearing a mask, or distinguishing real from synthetic speech. In this work, we introduce a new state-of-the-art paralinguistic representation derived fr…

Cited by 0SourceScholar
2018

Improved Lossy Image Compression With Priming and Spatially Adaptive Bit Rates for Recurrent Networks

CVPR 2018poster

We propose a method for lossy image compression based on recurrent, convolutional neural networks that outper- forms BPG (4:2:0), WebP, JPEG2000, and JPEG as mea- sured by MS-SSIM. We introduce three improvements over previous research that lead to this state-of-the-art result us- ing a single model…

Cited by 483SourcePDFScholar
2018

Style Tokens: Unsupervised Style Modeling, Control and Transfer in End-to-End Speech Synthesis

ICML 2018oral

In this work, we propose “global style tokens” (GSTs), a bank of embeddings that are jointly trained within Tacotron, a state-of-the-art end-to-end speech synthesis system. The embeddings are trained with no explicit labels, yet learn to model a large range of acoustic expressiveness. GSTs lead to a…

Cited by 1059SourcePDFScholar
2018

Towards End-to-End Prosody Transfer for Expressive Speech Synthesis with Tacotron

ICML 2018oral

We present an extension to the Tacotron speech synthesis architecture that learns a latent embedding space of prosody, derived from a reference acoustic representation containing the desired prosody. We show that conditioning Tacotron on this learned embedding space results in synthesized audio that…

Cited by 749SourcePDFScholar
2017

Full Resolution Image Compression With Recurrent Neural Networks

CVPR 2017oral

This paper presents a set of full-resolution lossy image compression methods based on neural networks. Each of the architectures we describe can provide variable compression rates during deployment without requiring retraining of the network: each network need only be trained once. All of our archit…

Cited by 1111PDFScholar