← Search

Roy Hirsch

5 accepted papers

2024

On the Semantic Latent Space of Diffusion-Based Text-To-Speech Models

ACL 2024short

The incorporation of Denoising Diffusion Models (DDMs) in the Text-to-Speech (TTS) domain is rising, providing great value in synthesizing high quality speech. Although they exhibit impressive audio quality, the extent of their semantic capabilities is unknown, and controlling their synthesized spee…

2024

Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

ICLR 2024poster

We present Spectron, a novel approach to adapting pre-trained large language models (LLMs) to perform spoken question answering (QA) and speech continuation. By endowing the LLM with a pre-trained speech encoder, our model becomes able to take speech inputs and generate speech outputs. The entire sy…

Cited by 40SourcePDFScholar
2023

Efficient Discovery and Effective Evaluation of Visual Perceptual Similarity: A Benchmark and Beyond

ICCV 2023poster

Visual similarities discovery (VSD) is an important task with broad e-commerce applications. Given an image of a certain object, the goal of VSD is to retrieve images of different objects with high perceptual visual similarity. Although being a highly addressed problem, the evaluation of proposed me…

Cited by 6PDFcodeScholar
2021

Cold Start Revisited: A Deep Hybrid Recommender with Cold-Warm Item Harmonization

ICASSP 2021accepted

Collaborative filtering-based recommender systems are known to suffer from the item cold-start problem. Most recent attempts to mitigate this problem presented parametric approaches, such as deep content based models. In this paper, we show that a straightforward application of parametric models may…

Cited by 0SourceScholar