← Search

Niko Moritz

18 accepted papers

2025

Directional Source Separation for Robust Speech Recognition on Smart Glasses

ICASSP 2025accepted

Modern smart glasses leverage machine learning to offer real-time transcriptions, considerably enriching human communication experiences. However, such systems frequently encounter challenges related to environmental noises, leading to decreased speech recognition. To improve voice quality, this wor…

Cited by 15SourceScholar
2025

M-BEST-RQ: A Multi-Channel Speech Foundation Model for Smart Glasses

ICASSP 2025accepted

The growing popularity of multi-channel wearable devices, such as smart glasses, has led to a surge of applications such as targeted speech recognition and enhanced hearing. However, current approaches to solve these tasks use independently trained models, which may not benefit from large amounts of…

Cited by 0SourceScholar
2025

Textless Streaming Speech-to-Speech Translation using Semantic Speech Tokens

ICASSP 2025accepted

Cascaded speech-to-speech translation systems often suffer from the error accumulation problem and high latency, which is a result of cascaded modules whose inference delays accumulate. In this paper, we propose a transducer-based speech translation model that outputs discrete speech tokens in a low…

Cited by 0SourceScholar
2025

Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition

ICASSP 2025accepted

We propose the joint speech translation and recognition (JSTAR) model that leverages the fast-slow cascaded encoder architecture for simultaneous end-to-end automatic speech recognition (ASR) and speech translation (ST). The model is transducer-based and uses a multi-objective training strategy that…

Cited by 0SourceScholar
2025

Transducer-Llama: Integrating LLMs into Streamable Transducer-based Speech Recognition

ICASSP 2025accepted

While large language models (LLMs) have been applied to automatic speech recognition (ASR), the task of making the model streamable remains a challenge. This paper proposes a novel model architecture, Transducer-Llama, that integrates LLMs into a Factorized Transducer (FT) model, naturally enabling…

Cited by 0SourceScholar
2024

AGADIR: Towards Array-Geometry Agnostic Directional Speech Recognition

ICASSP 2024accepted

Wearable devices like smart glasses are approaching the compute capability to seamlessly generate real-time closed captions for live conversations. We build on our recently introduced directional Automatic Speech Recognition (ASR) for smart glasses that have microphone arrays, which fuses multi-chan…

Cited by 0SourceScholar
2024

Effective Internal Language Model Training and Fusion for Factorized Transducer Model

ICASSP 2024accepted

The internal language model (ILM) of the neural transducer has been widely studied. In most prior work, it is mainly used for estimating the ILM score and is subsequently subtracted during inference to facilitate improved integration with external language models. Recently, various of factorized tra…

Cited by 0SourceScholar
2023

Anchored Speech Recognition with Neural Transducers

ICASSP 2023accepted

Neural transducers have achieved human level performance on standard speech recognition benchmarks. However, their performance significantly degrades in the presence of cross-talk, especially when the primary speaker has a low signal-to-noise ratio. Anchored speech recognition refers to a class of m…

Cited by 2SourceScholar
2023

SynthVSR: Scaling Up Visual Speech Recognition With Synthetic Supervision

CVPR 2023poster

Recently reported state-of-the-art results in visual speech recognition (VSR) often rely on increasingly large amounts of video data, while the publicly available transcribed video datasets are limited in size. In this paper, for the first time, we study the potential of leveraging synthetic visual…

Cited by 27SourcePDFScholar
2022

Advancing Momentum Pseudo-Labeling with Conformer and Initialization Strategy

ICASSP 2022accepted

Pseudo-labeling (PL), a semi-supervised learning (SSL) method where a seed model performs self-training using pseudo-labels generated from untranscribed speech, has been shown to enhance the performance of end-to-end automatic speech recognition (ASR). Our prior work proposed momentum pseudo-labelin…

Cited by 14SourceScholar
2022

Extended Graph Temporal Classification for Multi-Speaker End-to-End ASR

ICASSP 2022accepted

Graph-based temporal classification (GTC), a generalized form of the connectionist temporal classification loss, was recently proposed to improve automatic speech recognition (ASR) systems using graph-based supervision. For example, GTC was first used to encode an N-best list of pseudo-label sequenc…

Cited by 0SourceScholar
2021

Semi-Supervised Speech Recognition Via Graph-Based Temporal Classification

ICASSP 2021accepted

Semi-supervised learning has demonstrated promising results in automatic speech recognition (ASR) by self-training using a seed ASR model with pseudo-labels generated for unlabeled data. The effectiveness of this approach largely relies on the pseudo-label accuracy, for which typically only the 1-be…

Cited by 0SourceScholar
2021

Unsupervised Domain Adaptation for Speech Recognition via Uncertainty Driven Self-Training

ICASSP 2021accepted

The performance of automatic speech recognition (ASR) systems typically degrades significantly when the training and test data domains are mismatched. In this paper, we show that self-training (ST) combined with an uncertainty-based pseudo-label filtering approach can be effectively used for domain…

Cited by 0SourceScholar
2020

Unsupervised Speaker Adaptation Using Attention-Based Speaker Memory for End-to-End ASR

ICASSP 2020accepted

We propose an unsupervised speaker adaptation method inspired by the neural Turing machine for end-to-end (E2E) automatic speech recognition (ASR). The proposed model contains a memory block that holds speaker i-vectors extracted from the training data and reads relevant i-vectors from the memory th…

Cited by 0SourceScholar