← Search

Mohammad Zeineldeen

7 accepted papers

2025

The Conformer Encoder May Reverse the Time Dimension

ICASSP 2025accepted

We sometimes observe monotonically decreasing cross-attention weights in our Conformer-based global attention-based encoder-decoder (AED) models, negatively affecting performance compared to monotonically increasing attention weights. Further investigation shows that the Conformer encoder reverses t…

Cited by 2SourceScholar
2024

Chunked Attention-Based Encoder-Decoder Model for Streaming Speech Recognition

ICASSP 2024accepted

We study a streamable attention-based encoder-decoder model in which either the decoder, or both the encoder and decoder, operate on pre-defined, fixed-size windows called chunks. A special end-of-chunk (EOC) symbol advances from one chunk to the next chunk, effectively replacing the conventional en…

Cited by 0SourceScholar
2023

Enhancing and Adversarial: Improve ASR with Speaker Labels

ICASSP 2023accepted

ASR can be improved by multi-task learning (MTL) with domain enhancing or domain adversarial training, which are two opposite objectives with the aim to increase/decrease domain variance towards domain-aware/agnostic ASR, respectively. In this work, we study how to best apply these two opposite obje…

Cited by 0SourceScholar
2023

Improving Language Model Integration for Neural Machine Translation

ACL 2023findings

The integration of language models for neural machine translation has been extensively studied in the past. It has been shown that an external language model, trained on additional target-side monolingual data, can help improve translation quality. However, there has always been the assumption that…

Cited by 4SourcePDFScholar
2023

Robust Knowledge Distillation from RNN-T Models with Noisy Training Labels Using Full-Sum Loss

ICASSP 2023accepted

This work studies knowledge distillation (KD) and addresses its constraints for recurrent neural network transducer (RNN-T) models. In hard distillation, a teacher model transcribes large amounts of unlabelled speech to train a student model. Soft distillation is another popular KD method that disti…

Cited by 0SourceScholar
2022

Conformer-Based Hybrid ASR System For Switchboard Dataset

ICASSP 2022accepted

The recently proposed conformer architecture has been successfully used for end-to-end automatic speech recognition (ASR) architectures achieving state-of-the-art performance on different datasets. To our best knowledge, the impact of using conformer acoustic model for hybrid ASR is not investigated…

Cited by 0SourceScholar
2020

Layer-Normalized LSTM for Hybrid-Hmm and End-To-End ASR

ICASSP 2020accepted

Training deep neural networks is often challenging in terms of training stability. It often requires careful hyperparameter tuning or a pretraining scheme to converge. Layer normalization (LN) has shown to be a crucial ingredient in training deep encoder-decoder models. We explore various LN long sh…

Cited by 0SourceScholar