← Search

Zvi Kons

4 accepted papers

2024

Speak While You Think: Streaming Speech Synthesis During Text Generation

ICASSP 2024accepted

Large Language Models (LLMs) demonstrate impressive capabilities, yet interaction with these models is mostly facilitated through text. Using Text-To-Speech to synthesize LLM outputs typically results in notable latency, which is impractical for fluent voice conversations. We propose LLM2Speech, an…

Cited by 0SourceScholar
2022

A New Data Augmentation Method for Intent Classification Enhancement and its Application on Spoken Conversation Datasets

ICASSP 2022accepted

Intent classifiers are vital to the successful operation of virtual agent systems. This is especially so in voice activated systems where the data can be noisy with many ambiguous directions for user intents. Before operation begins, these classifiers are generally lacking in real-world training dat…

Cited by 0SourceScholar
2021

RNN Transducer Models for Spoken Language Understanding

ICASSP 2021accepted

We present a comprehensive study on building and adapting RNN transducer (RNN-T) models for spoken language understanding (SLU). These end-to-end (E2E) models are constructed in three practical settings: a case where verbatim transcripts are available, a constrained case where the only available ann…

Cited by 0SourceScholar
2020

Leveraging Unpaired Text Data for Training End-To-End Speech-to-Intent Systems

ICASSP 2020accepted

Training an end-to-end (E2E) neural network speech-to-intent (S2I) system that directly extracts intents from speech requires large amounts of intent-labeled speech data, which is time consuming and expensive to collect. Initializing the S2I model with an ASR model trained on copious speech data can…

Cited by 0SourceScholar