← Search

Matthew Wiesner

12 accepted papers

2026

SE-DiCoW: Self-Enrolled Diarization-Conditioned Whisper

ICASSP 2026oral

Speaker-attributed automatic speech recognition (ASR) in multi-speaker environments remains a major challenge. While some approaches achieve strong performance when fine-tuned on specific domains, few systems generalize well across out-of-domain datasets. Our prior work, Diarization-Conditioned Whis…

Cited by 4SourcePDFScholar
2025

HLTCOE Submission to the VoicePrivacy Attacker Challenge

ICASSP 2025accepted

We describe our submission to the 2024 VoicePrivacy Attacker Challenge. We propose three main categories of methods to improve ASV performance against anonymized speech: improvements to the underlying classifier, alternative distance metrics when computing ASV scores, and kNN-VC normalization. By si…

Cited by 0SourceScholar
2025

Target Speaker ASR with Whisper

ICASSP 2025accepted

We propose a novel approach to enable the use of large, single-speaker ASR models, such as Whisper, for target speaker ASR. The key claim of this method is that it is much easier to model relative differences among speakers by learning to condition on frame-level diarization outputs than to learn th…

Cited by 0SourceScholar
2025

Whisper-UT: A Unified Translation Framework for Speech and Text

EMNLP 2025

Encoder-decoder models have achieved remarkable success in speech and text tasks, yet efficiently adapting these models to diverse uni/multi-modal scenarios remains an open challenge. In this paper, we propose Whisper-UT, a unified and efficient framework that leverages lightweight adapters to enabl

Cited by 0SourcePDFScholar
2024

Less Peaky and More Accurate CTC Forced Alignment by Label Priors

ICASSP 2024accepted

Connectionist temporal classification (CTC) models are known to have peaky output distributions. Such behavior is not a problem for automatic speech recognition (ASR), but it can cause inaccurate forced alignments (FA), especially at finer granularity, e.g., phoneme level. This paper aims at allevia…

Cited by 0SourceScholar
2024

Speech Collage: Code-Switched Audio Generation by Collaging Monolingual Corpora

ICASSP 2024accepted

Designing effective automatic speech recognition (ASR) systems for Code-Switching (CS) often depends on the availability of the transcribed CS resources. To address data scarcity, this paper introduces Speech Collage, a method that synthesizes CS data from monolingual corpora by splicing audio segme…

Cited by 0SourceScholar
2024

Where are you from? Geolocating Speech and Applications to Language Identification

NAACL 2024long

We train models to answer the question, Where are you from? and show how such models can be repurposed for language identification (LID). To our knowledge, this paper is the first to introduce data sources, methods and models to tackle the task of geolocation of speech at a global scale, and the fir…

Cited by 1SourcePDFScholar
2023

Building Keyword Search System from End-To-End Asr Systems

ICASSP 2023accepted

Keyword search (KWS) systems are commonly built on top of existing automatic speech recognition (ASR) systems. However, end-to-end (E2E) ASR models are not naturally equipped with word-level timing information or confidence. Existing methods for re-purposing E2E ASR systems for KWS are largely heuri…

Cited by 0SourceScholar
2023

Towards Zero-Shot Code-Switched Speech Recognition

ICASSP 2023accepted

In this work, we seek to build effective code-switched (CS) automatic speech recognition systems (ASR) under the zero-shot set-ting where no transcribed CS speech data is available for training. Previously proposed frameworks which conditionally factorize the bilingual task into its constituent mono…

Cited by 0SourceScholar
2022

Injecting Text and Cross-Lingual Supervision in Few-Shot Learning from Self-Supervised Models

ICASSP 2022accepted

Self-supervised model pretraining has recently garnered significant interest. However, using additional resources in fine-tuning these models has received less attention. We demonstrate how universal phoneset acoustic models can leverage cross-lingual supervision to improve transfer of pretrained se…

Cited by 0SourceScholar
2021

End-to-end ASR to jointly predict transcriptions and linguistic annotations

NAACL 2021long

We propose a Transformer-based sequence-to-sequence model for automatic speech recognition (ASR) capable of simultaneously transcribing and annotating audio with linguistic information such as phonemic transcripts or part-of-speech (POS) tags. Since linguistic information is important in natural lan…

Cited by 13SourcePDFScholar
2015

Towards machines that know when they do not know: Summary of work done at 2014 Frederick Jelinek Memorial Workshop

ICASSP 2015accepted

A group of junior and senior researchers gathered as a part of the 2014 Frederick Jelinek Memorial Workshop in Prague to address the problem of predicting the accuracy of a nonlinear Deep Neural Network probability estimator for unknown data in a different application domain from the domain in which…

Cited by 0SourceScholar