← Search

Albert Zeyer

10 accepted papers

2025

The Conformer Encoder May Reverse the Time Dimension

ICASSP 2025accepted

We sometimes observe monotonically decreasing cross-attention weights in our Conformer-based global attention-based encoder-decoder (AED) models, negatively affecting performance compared to monotonically increasing attention weights. Further investigation shows that the Conformer encoder reverses t…

Cited by 0SourceScholar
2024

Chunked Attention-Based Encoder-Decoder Model for Streaming Speech Recognition

ICASSP 2024accepted

We study a streamable attention-based encoder-decoder model in which either the decoder, or both the encoder and decoder, operate on pre-defined, fixed-size windows called chunks. A special end-of-chunk (EOC) symbol advances from one chunk to the next chunk, effectively replacing the conventional en…

Cited by 0SourceScholar
2020

A Comprehensive Study of Residual CNNS for Acoustic Modeling in ASR

ICASSP 2020accepted

Long short-term memory (LSTM) networks are the dominant architecture for large vocabulary continuous speech recognition (LVCSR) acoustic modeling due to their good performance. However, LSTMs are hard to tune and computationally expensive. To build a system with lower computational costs and which a…

Cited by 0SourceScholar
2020

Exploring A Zero-Order Direct Hmm Based on Latent Attention for Automatic Speech Recognition

ICASSP 2020accepted

In this paper, we study a simple yet elegant latent variable attention model for automatic speech recognition (ASR) which enables an integration of attention sequence modeling into the direct hidden Markov model (HMM) concept. We use a sequence of hidden variables that establishes a mapping from out…

Cited by 0SourceScholar
2020

Generating Synthetic Audio Data for Attention-Based Speech Recognition Systems

ICASSP 2020accepted

Recent advances in text-to-speech (TTS) led to the development of flexible multi-speaker end-to-end TTS systems. We extend state-of-the-art attention-based automatic speech recognition (ASR) systems with synthetic audio generated by a TTS system trained only on the ASR corpora itself. ASR and TTS sy…

Cited by 0SourceScholar
2020

Layer-Normalized LSTM for Hybrid-Hmm and End-To-End ASR

ICASSP 2020accepted

Training deep neural networks is often challenging in terms of training stability. It often requires careful hyperparameter tuning or a pretraining scheme to converge. Layer normalization (LN) has shown to be a crucial ingredient in training deep encoder-decoder models. We explore various LN long sh…

Cited by 0SourceScholar
2019

On Using 2D Sequence-to-sequence Models for Speech Recognition

ICASSP 2019accepted

Attention-based sequence-to-sequence models have shown promising results in automatic speech recognition. Using these architectures, one-dimensional input and output sequences are related by an attention approach, thereby replacing more explicit alignment processes, like in classical HMM-based model…

Cited by 0SourceScholar
2017

A comprehensive study of deep bidirectional LSTM RNNS for acoustic modeling in speech recognition

ICASSP 2017accepted

Recent experiments show that deep bidirectional long short-term memory (BLSTM) recurrent neural network acoustic models outperform feedforward neural networks for automatic speech recognition (ASR). However, their training requires a lot of tuning and experience. In this work, we provide a comprehen…

Cited by 0SourceScholar
2017

Returnn: The RWTH extensible training framework for universal recurrent neural networks

ICASSP 2017accepted

In this work we release our extensible and easily configurable neural network training software. It provides a rich set of functional layers with a particular focus on efficient training of recurrent neural network topologies on multiple GPUs. The source of the software package is public and freely…

Cited by 0SourceScholar