← Search

Joseph Keshet

13 accepted papers

2026

Beyond Transcription: Mechanistic Interpretability in ASR

AAAI 2026technical

Interpretability methods have recently gained significant attention, particularly in the context of large language models, enabling insights into linguistic representations, error detection, and model behaviors such as hallucinations and repetitions. However, these techniques remain underexplored in

Cited by 0SourcePDFScholar
2026

Joint Enhancement and Classification using Coupled Diffusion Models of Signals and Logits

ICML 2026poster

Robust classification in noisy environments remains a fundamental challenge in machine learning. Standard approaches typically treat signal enhancement and classification as separate, sequential stages: first enhancing the signal and then applying a classifier. This approach fails to leverage the se…

Cited by 0SourceScholar
2025

Whisper in Medusa's Ear: Multi-head Efficient Decoding for Transformer-based ASR

ICASSP 2025accepted

Large transformer-based models have significant potential for speech transcription and translation. Their self-attention mechanisms and parallel processing enable them to capture complex patterns and dependencies in audio sequences. However, this potential comes with challenges, as these large and c…

Cited by 0SourceScholar
2024

DiffAR: Denoising Diffusion Autoregressive Model for Raw Speech Waveform Generation

ICLR 2024poster

Diffusion models have recently been shown to be relevant for high-quality speech generation. Most work has been focused on generating spectrograms, and as such, they further require a subsequent model to convert the spectrogram to a waveform (i.e., a vocoder). This work proposes a diffusion probabil…

2024

Open-Vocabulary Keyword-Spotting with Adaptive Instance Normalization

ICASSP 2024accepted

Open vocabulary keyword spotting is a crucial and challenging task in automatic speech recognition (ASR) that focuses on detecting user-defined keywords within a spoken utterance. Keyword spotting methods commonly map the audio utterance and keyword into a joint embedding space to obtain some affini…

Cited by 0SourceScholar
2021

CNN-Based Spoken Term Detection and Localization without Dynamic Programming

ICASSP 2021accepted

In this paper, we propose a spoken term detection algorithm for simultaneous prediction and localization of in-vocabulary and out-of-vocabulary terms within an audio segment. The proposed algorithm infers whether a term was uttered within a given speech signal or not by predicting the word embedding…

Cited by 0SourceScholar
2018

Fooling End-To-End Speaker Verification With Adversarial Examples

ICASSP 2018accepted

Automatic speaker verification systems are increasingly used as the primary means to authenticate costumers. Recently, it has been proposed to train speaker verification systems using end-to-end deep neural models. In this paper, we show that such systems are vulnerable to adversarial example attack…

Cited by 0SourceScholar
2018

Out-of-Distribution Detection using Multiple Semantic Label Representations

NeurIPS 2018poster

Deep Neural Networks are powerful models that attained remarkable results on a variety of tasks. These models are shown to be extremely efficient when training and test data are drawn from the same distribution. However, it is not clear how a network will act when it is fed with an out-of-distributi…

2017

Houdini: Fooling Deep Structured Visual and Speech Recognition Models with Adversarial Examples

NeurIPS 2017poster

Generating adversarial examples is a critical step for evaluating and improving the robustness of learning machines. So far, most existing methods only work for classification and are not designed to alter the true performance measure of the problem at hand. We introduce a novel flexible approach na…

Cited by 226SourcePDFScholar
2017

Sequence segmentation using joint RNN and structured prediction models

ICASSP 2017accepted

We describe and analyze a simple and effective algorithm for sequence segmentation applied to speech processing tasks. We propose a neural architecture that is composed of two modules trained jointly: a recurrent neural network (RNN) module and a structured prediction model. The RNN outputs are cons…

Cited by 0SourceScholar
2016

The relationship of voice onset time and Voice Offset Time to physical age

ICASSP 2016accepted

In a speech signal, Voice Onset Time (VOT) is the period between the release of a plosive and the onset of vocal cord vibrations in the production of the following sound. Voice Offset Time (VOFT), on the other hand, is the period between the end of a voiced sound and the release of the following plo…

Cited by 0SourceScholar