← Search

Ron Hoory

10 accepted papers

2024

Speak While You Think: Streaming Speech Synthesis During Text Generation

ICASSP 2024accepted

Large Language Models (LLMs) demonstrate impressive capabilities, yet interaction with these models is mostly facilitated through text. Using Text-To-Speech to synthesize LLM outputs typically results in notable latency, which is impractical for fluent voice conversations. We propose LLM2Speech, an…

Cited by 0SourceScholar
2023

Modeling Turn-Taking in Human-To-Human Spoken Dialogue Datasets Using Self-Supervised Features

ICASSP 2023accepted

Self-supervised pre-trained models have consistently delivered state-of-art results in the fields of natural language and speech processing. However, we argue that their merits for modeling Turn-Taking for spoken dialogue systems still need further investigation. Due to that, in this paper we intro-…

Cited by 0SourceScholar
2022

A New Data Augmentation Method for Intent Classification Enhancement and its Application on Spoken Conversation Datasets

ICASSP 2022accepted

Intent classifiers are vital to the successful operation of virtual agent systems. This is especially so in voice activated systems where the data can be noisy with many ambiguous directions for user intents. Before operation begins, these classifiers are generally lacking in real-world training dat…

Cited by 0SourceScholar
2022

Speaker Normalization for Self-Supervised Speech Emotion Recognition

ICASSP 2022accepted

Large speech emotion recognition datasets are hard to obtain, and small datasets may contain biases. Deep-net-based classifiers, in turn, are prone to exploit those biases and find shortcuts such as speaker characteristics. These shortcuts usually harm a model’s ability to generalize. To address thi…

Cited by 64SourceScholar
2022

Speech Emotion Recognition Using Self-Supervised Features

ICASSP 2022accepted

Self-supervised pre-trained features have consistently delivered state-of-art results in the field of natural language processing (NLP); however, their merits in the field of speech emotion recognition (SER) still need further investigation. In this paper we introduce a modular End-to-End (E2E) SER…

Cited by 0SourceScholar
2022

Towards A Common Speech Analysis Engine

ICASSP 2022accepted

Recent innovations in self-supervised representation learning have led to remarkable advances in natural language processing. That said, in the speech processing domain, self-supervised representation learning-based systems are not yet considered state-of-the-art.We propose leveraging recent advance…

Cited by 0SourceScholar
2021

RNN Transducer Models for Spoken Language Understanding

ICASSP 2021accepted

We present a comprehensive study on building and adapting RNN transducer (RNN-T) models for spoken language understanding (SLU). These end-to-end (E2E) models are constructed in three practical settings: a case where verbatim transcripts are available, a constrained case where the only available ann…

Cited by 0SourceScholar
2020

Leveraging Unpaired Text Data for Training End-To-End Speech-to-Intent Systems

ICASSP 2020accepted

Training an end-to-end (E2E) neural network speech-to-intent (S2I) system that directly extracts intents from speech requires large amounts of intent-labeled speech data, which is time consuming and expensive to collect. Initializing the S2I model with an ASR model trained on copious speech data can…

Cited by 0SourceScholar
2017

Voice-transformation-based data augmentation for prosodic classification

ICASSP 2017accepted

In this work we explore data-augmentation techniques for the task of improving the performance of a supervised recurrent-neural-network classifier tasked with predicting prosodic-boundary and pitch-accent labels. The technique is based on applying voice transformations to the training data that modi…

Cited by 12SourceScholar
2016

Using continuous lexical embeddings to improve symbolic-prosody prediction in a text-to-speech front-end

ICASSP 2016accepted

The prediction of symbolic prosodic categories from text is an important, but challenging, natural-language processing task given the various ways in which an input can be realized, and the fact that knowledge about what features determine this realization is incomplete or inaccessible to the model.…

Cited by 0SourceScholar