← Search

Hagai Aronowitz

12 accepted papers

2025

A Non-autoregressive Model for Joint STT and TTS

ICASSP 2025accepted

In this paper, we take a step towards jointly modeling automatic speech recognition (STT) and speech synthesis (TTS) in a fully non-autoregressive way. We develop a novel multimodal framework capable of handling the speech and text modalities as input either individually or together. The proposed mo…

Cited by 0SourceScholar
2023

Modeling Turn-Taking in Human-To-Human Spoken Dialogue Datasets Using Self-Supervised Features

ICASSP 2023accepted

Self-supervised pre-trained models have consistently delivered state-of-art results in the fields of natural language and speech processing. However, we argue that their merits for modeling Turn-Taking for spoken dialogue systems still need further investigation. Due to that, in this paper we intro-…

Cited by 0SourceScholar
2022

Speaker Normalization for Self-Supervised Speech Emotion Recognition

ICASSP 2022accepted

Large speech emotion recognition datasets are hard to obtain, and small datasets may contain biases. Deep-net-based classifiers, in turn, are prone to exploit those biases and find shortcuts such as speaker characteristics. These shortcuts usually harm a model’s ability to generalize. To address thi…

Cited by 64SourceScholar
2022

Speech Emotion Recognition Using Self-Supervised Features

ICASSP 2022accepted

Self-supervised pre-trained features have consistently delivered state-of-art results in the field of natural language processing (NLP); however, their merits in the field of speech emotion recognition (SER) still need further investigation. In this paper we introduce a modular End-to-End (E2E) SER…

Cited by 0SourceScholar
2022

Towards A Common Speech Analysis Engine

ICASSP 2022accepted

Recent innovations in self-supervised representation learning have led to remarkable advances in natural language processing. That said, in the speech processing domain, self-supervised representation learning-based systems are not yet considered state-of-the-art.We propose leveraging recent advance…

Cited by 0SourceScholar
2018

Robust Audiovisual Liveness Detection for Biometric Authentication Using Deep Joint Embedding and Dynamic Time Warping

ICASSP 2018accepted

We address the problem of liveness detection in audiovisual recordings for preventing spoofing attacks in biometric authentication systems. We assume that liveness is detected from a recording of a speaker saying a predefined phrase and that another recording of the same phrase is a priori available…

Cited by 0SourceScholar
2016

Audio enhancing with DNN autoencoder for speaker recognition

ICASSP 2016accepted

In this paper we present a design of a DNN-based autoencoder for speech enhancement and its use for speaker recognition systems for distant microphones and noisy data. We started with augmenting the Fisher database with artificially noised and reverberated data and trained the autoencoder to map noi…

Cited by 0SourceScholar
2016

Speaker recognition using matched filters

ICASSP 2016accepted

Nowadays state-of-the-art speaker recognition systems obtain quite accurate results for both text-independent and text-dependent tasks as long as they are trained on a fair amount of development data from the target domain, and as long as the target data is clean. In this work we investigate the use…

Cited by 0SourceScholar