← Search

Stefan Braun

5 accepted papers

2023

Neural Transducer Training: Reduced Memory Consumption with Sample-Wise Computation

ICASSP 2023accepted

The neural transducer is an end-to-end model for automatic speech recognition (ASR). While the model is well-suited for streaming ASR, the training process remains challenging. During training, the memory requirements may quickly exceed the capacity of state-of-the-art GPUs, limiting batch size and…

Cited by 0SourceScholar
2023

Variable Attention Masking for Configurable Transformer Transducer Speech Recognition

ICASSP 2023accepted

This work studies the use of attention masking in transformer transducer based speech recognition for building a single configurable model for different deployment scenarios. We present a comprehensive set of experiments comparing fixed masking, where the same attention mask is applied at every fram…

Cited by 0SourceScholar
2021

SapAugment: Learning A Sample Adaptive Policy for Data Augmentation

ICASSP 2021accepted

Data augmentation methods usually apply the same augmentation (or a mix of them) to all the training samples. For example, to perturb data with noise, the noise is sampled from a Normal distribution with a fixed standard deviation, for all samples. We hypothesize that a hard sample with high trainin…

Cited by 0SourceScholar
2019

Event-driven Pipeline for Low-latency Low-compute Keyword Spotting and Speaker Verification System

ICASSP 2019accepted

This work presents an event-driven acoustic sensor processing pipeline to power a low-resource voice-activated smart assistant. The pipeline includes four major steps; namely localization, source separation, keyword spotting (KWS) and speaker verification (SV). The pipeline is driven by a front-end…

Cited by 0SourceScholar