← Search

Björn Hoffmeister

10 accepted papers

2024

An Efficient Self-Learning Framework For Interactive Spoken Dialog Systems

ICML 2024poster

Dialog systems, such as voice assistants, are expected to engage with users in complex, evolving conversations. Unfortunately, traditional automatic speech recognition (ASR) systems deployed in such applications are usually trained to recognize each turn independently and lack the ability to adapt t…

Cited by 0SourcePDFScholar
2024

Task Oriented Dialogue as a Catalyst for Self-Supervised Automatic Speech Recognition

ICASSP 2024accepted

While word error rates of automatic speech recognition (ASR) systems have consistently fallen, natural language understanding (NLU) applications built on top of ASR systems still attribute significant numbers of failures to low-quality speech recognition results. Existing assistant systems collect l…

Cited by 0SourceScholar
2023

Domain Adaptation with External Off-Policy Acoustic Catalogs for Scalable Contextual End-to-End Automated Speech Recognition

ICASSP 2023accepted

Despite improvements to the generalization performance of automated speech recognition (ASR) models, specializing ASR models for downstream tasks remains a challenging task, primarily due to reduced data availability (necessitating increased data collection), and rapidly shifting data distributions…

Cited by 0SourceScholar
2022

Multi-Modal Pre-Training for Automated Speech Recognition

ICASSP 2022accepted

Traditionally, research in automated speech recognition has focused on local-first encoding of audio representations to predict the spoken phonemes in an utterance. Unfortunately, approaches relying on such hyper-local information tend to be vulnerable to both local-level corruption (such as audio-f…

Cited by 0SourceScholar
2019

End-to-end Anchored Speech Recognition

ICASSP 2019accepted

Voice-controlled house-hold devices, like Amazon Echo or Google Home, face the problem of performing speech recognition of device-directed speech in the presence of interfering background speech, i.e., background noise and interfering speech from another person or media device in proximity need to b…

Cited by 0SourceScholar
2019

Frequency Domain Multi-channel Acoustic Modeling for Distant Speech Recognition

ICASSP 2019accepted

Conventional far-field automatic speech recognition (ASR) systems typically employ microphone array techniques for speech enhancement in order to improve robustness against noise or reverberation. However, such speech enhancement techniques do not always yield ASR accuracy improvement because the op…

Cited by 0SourceScholar
2019

Improving Noise Robustness of Automatic Speech Recognition via Parallel Data and Teacher-student Learning

ICASSP 2019accepted

For real-world speech recognition applications, noise robustness is still a challenge. In this work, we adopt the teacher-student (T/S) learning technique using a parallel clean and noisy corpus for improving automatic speech recognition (ASR) performance under multimedia noise. On top of that, we a…

Cited by 0SourceScholar
2019

Multi-geometry Spatial Acoustic Modeling for Distant Speech Recognition

ICASSP 2019accepted

The use of spatial information with multiple microphones can improve far-field automatic speech recognition (ASR) accuracy. However, conventional microphone array techniques degrade speech enhancement performance when there is an array geometry mismatch between design and test conditions. Moreover,…

Cited by 0SourceScholar
2018

Combining Acoustic Embeddings and Decoding Features for End-of-Utterance Detection in Real-Time Far-Field Speech Recognition Systems

ICASSP 2018accepted

We present an end-of-utterance detector for real-time automatic speech recognition in far-field scenarios. The proposed system consists of three components: a long short-term memory (LSTM) neural network trained on acoustic features, an LSTM trained on l-best recognition hypotheses of the automatic…

Cited by 0SourceScholar
2018

Monophone-Based Background Modeling for Two-Stage On-Device Wake Word Detection

ICASSP 2018accepted

Accurate on-device wake word detection is crucial to products with far-field voice control such as the Amazon Echo. It is quite challenging to build a wake word system with both low False Reject Rate (FRR) and low False Alarm Rate (FAR) in real scenarios where there are various types of background s…

Cited by 0SourceScholar