← Search

Andreas Schwarz

4 accepted papers

2025

SIFT-50M: A Large-Scale Multilingual Dataset for Speech Instruction Fine-Tuning

ACL 2025long

We introduce SIFT (Speech Instruction Fine-Tuning), a 50M-example dataset designed for instruction fine-tuning and pre-training of speech-text large language models (LLMs). SIFT-50M is built from publicly available speech corpora, which collectively contain 14K hours of speech, and leverages LLMs al…

2024

Promptformer: Prompted Conformer Transducer for ASR

ICASSP 2024accepted

Context cues carry information which can improve multi-turn interactions in automatic speech recognition (ASR) systems. In this paper, we introduce a novel mechanism inspired by hyper-prompting to fuse textual context with acoustic representations in the attention mechanism. Results on a test set wi…

Cited by 0SourceScholar
2016

A new uncertainty decoding scheme for DNN-HMM hybrid systems with multichannel speech enhancement

ICASSP 2016accepted

Uncertainty decoding combines a probabilistic feature description with the acoustic model of a speech recognition system. For DNN-HMM hybrid systems, this can be realized by averaging the DNN outputs produced by a finite set of feature samples (drawn from an estimated probability distribution). In t…

Cited by 0SourceScholar
2015

Spatial diffuseness features for DNN-based speech recognition in noisy and reverberant environments

ICASSP 2015accepted

We propose a spatial diffuseness feature for deep neural network (DNN)-based automatic speech recognition to improve recognition accuracy in reverberant and noisy environments. The feature is computed in real-time from multiple microphone signals without requiring knowledge or estimation of the dire…

Cited by 0SourceScholar