← Search

Zoltán Tüske

12 accepted papers

2024

SADA: Saudi Audio Dataset for Arabic

ICASSP 2024accepted

Arabic is among the most challenging languages in the world. Unfortunately, the scarcity of Arabic datasets makes studies in Arabic speech technology demanding. This paper introduces SADA, the Saudi Audio Dataset for Arabic, with 668 hours of high-quality audio suitable for supervised training. The…

Cited by 0SourceScholar
2022

Improving End-to-end Models for Set Prediction in Spoken Language Understanding

ICASSP 2022accepted

The goal of spoken language understanding (SLU) systems is to determine the meaning of the input speech signal, unlike speech recognition which aims to produce verbatim transcripts. Advances in end-to-end (E2E) speech modeling have made it possible to train solely on semantic entities, which are far…

Cited by 0SourceScholar
2021

Advancing RNN Transducer Technology for Speech Recognition

ICASSP 2021accepted

We investigate a set of techniques for RNN Transducers (RNN-Ts) that were instrumental in lowering the word error rate on three different tasks (Switchboard 300 hours, conversational Spanish 780 hours and conversational Italian 900 hours). The techniques pertain to architectural changes, speaker ada…

Cited by 0SourceScholar
2021

End-to-End Spoken Language Understanding Using Transformer Networks and Self-Supervised Pre-Trained Features

ICASSP 2021accepted

Transformer networks and self-supervised pre-training have consistently delivered state-of-art results in the field of natural language processing (NLP); however, their merits in the field of spoken language understanding (SLU) still need further investigation. In this paper we introduce a modular E…

Cited by 0SourceScholar
2021

RNN Transducer Models for Spoken Language Understanding

ICASSP 2021accepted

We present a comprehensive study on building and adapting RNN transducer (RNN-T) models for spoken language understanding (SLU). These end-to-end (E2E) models are constructed in three practical settings: a case where verbatim transcripts are available, a constrained case where the only available ann…

Cited by 0SourceScholar
2019

English Broadcast News Speech Recognition by Humans and Machines

ICASSP 2019accepted

With recent advances in deep learning, considerable attention has been given to achieving automatic speech recognition performance close to human performance on tasks like conversational telephone speech (CTS) recognition. In this paper we evaluate the usefulness of these proposed techniques on broa…

Cited by 0SourceScholar
2019

Sequence Noise Injected Training for End-to-end Speech Recognition

ICASSP 2019accepted

We present a simple noise injection algorithm for training end-to-end ASR models which consists in adding to the spectra of training utterances the scaled spectra of random utterances of comparable length. We conjecture that the sequence information of the "noise" utterances is important and verify…

Cited by 0SourceScholar
2018

Acoustic Modeling of Speech Waveform Based on Multi-Resolution, Neural Network Signal Processing

ICASSP 2018accepted

Recently, several papers have demonstrated that neural networks (NN) are able to perform the feature extraction as part of the acoustic model. Motivated by the Gammatone feature extraction pipeline, in this paper we extend the waveform based NN model by a second level of time-convolutional element.…

Cited by 0SourceScholar
2018

Evolutionary Stochastic Gradient Descent for Optimization of Deep Neural Networks

NeurIPS 2018poster

We propose a population-based Evolutionary Stochastic Gradient Descent (ESGD) framework for optimizing deep neural networks. ESGD combines SGD and gradient-free evolutionary algorithms as complementary algorithms in one framework in which the optimization alternates between the SGD step and evolutio…

2016

Investigation on log-linear interpolation of multi-domain neural network language model

ICASSP 2016accepted

Inspired by the success of multi-task training in acoustic modeling, this paper investigates a new architecture for a multi-domain neural network based language model (NNLM). The proposed model has several shared hidden layers and domain-specific output layers. As will be shown, the log-linear inter…

Cited by 0SourceScholar
2015

Integrating Gaussian mixtures into deep neural networks: Softmax layer with hidden variables

ICASSP 2015accepted

In the hybrid approach, neural network output directly serves as hidden Markov model (HMM) state posterior probability estimates. In contrast to this, in the tandem approach neural network output is used as input features to improve classic Gaussian mixture model (GMM) based emission probability est…

Cited by 45SourceScholar