← Search

John R. Hershey

33 accepted papers

2025

I-Con: A Unifying Framework for Representation Learning

ICLR 2025poster

As the field of representation learning grows, there has been a proliferation of different loss functions to solve different classes of problems. We introduce a single information-theoretic equation that generalizes a large collection of mod- ern loss functions in machine learning. In particular, we…

Cited by 0SourcePDFScholar
2025

Towards Sub-millisecond Latency Real-Time Speech Enhancement Models on Hearables

ICASSP 2025accepted

Low latency models are critical for real-time speech enhancement applications, such as hearing aids and hearables. However, the sub-millisecond latency space for resource-constrained hearables remains underexplored. We demonstrate speech enhancement using a computationally efficient minimum-phase FI…

Cited by 0SourceScholar
2024

Separating the "Chirp" from the "Chat": Self-supervised Visual Grounding of Sound and Language

CVPR 2024poster

We present DenseAV a novel dual encoder grounding architecture that learns high-resolution semantically meaningful and audio-visual aligned features solely through watching videos. We show that DenseAV can discover the "meaning" of words and the "location" of sounds without explicit localization sup…

2022

Adapting Speech Separation to Real-World Meetings using Mixture Invariant Training

ICASSP 2022accepted

The recently-proposed mixture invariant training (MixIT) is an unsupervised method for training single-channel sound separation models because it does not require ground-truth isolated reference sources. In this paper, we investigate using MixIT to adapt a separation model on real far-field overlapp…

Cited by 0SourceScholar
2022

AudioScopeV2: Audio-Visual Attention Architectures for Calibrated Open-Domain On-Screen Sound Separation

ECCV 2022poster

"We introduce AudioScopeV2, a state-of-the-art universal audio-visual on-screen sound separation system which is capable of learning to separate sounds and associate them with on-screen objects by looking at in-the-wild videos. We identify several limitations of previous work on audio-visual on-scre…

2021

End-To-End Diarization for Variable Number of Speakers with Local-Global Networks and Discriminative Speaker Embeddings

ICASSP 2021accepted

We present an end-to-end deep network model that performs meeting diarization from single-channel audio recordings. End-to-end diarization models have the advantage of handling speaker overlap and enabling straightforward handling of discriminative training, unlike traditional clustering-based diari…

Cited by 0SourceScholar
2021

Into the Wild with AudioScope: Unsupervised Audio-Visual Separation of On-Screen Sounds

ICLR 2021poster

Recent progress in deep learning has enabled many advances in sound separation and visual scene understanding. However, extracting sound sources which are apparent in natural videos remains an open problem. In this work, we present AudioScope, a novel audio-visual sound separation framework that can…

Cited by 86SourcePDFScholar
2021

Sound Event Detection and Separation: A Benchmark on Desed Synthetic Soundscapes

ICASSP 2021accepted

We propose a benchmark of state-of-the-art sound event detection systems (SED). We design synthetic evaluation sets to focus on specific sound event detection challenges. We analyze the performance of the submissions to DCASE 2020 Task 4 as a function of time-related modifications (time position of…

Cited by 0SourceScholar
2021

What's all the Fuss about Free Universal Sound Separation Data?

ICASSP 2021accepted

We introduce the Free Universal Sound Separation (FUSS) dataset, a new corpus for experiments in separating mixtures of an unknown number of sounds from an open domain of sound types. The dataset consists of 23 hours of single-source audio data drawn from 357 classes, which are used to create mixtur…

Cited by 0SourceScholar
2020

Improving Universal Sound Separation Using Sound Classification

ICASSP 2020accepted

Deep learning approaches have recently achieved impressive performance on both audio source separation and sound classification. Most audio source separation approaches focus only on separating sources belonging to a restricted domain of source classes, such as speech and music. However, recent work…

Cited by 0SourceScholar
2020

Unsupervised Sound Separation Using Mixture Invariant Training

NeurIPS 2020spotlight

In recent years, rapid progress has been made on the problem of single-channel sound separation using supervised training of deep neural networks. In such supervised approaches, a model is trained to predict the component sources from synthetic mixtures created by adding up isolated ground-truth sou…

2019

Differentiable Consistency Constraints for Improved Deep Speech Enhancement

ICASSP 2019accepted

In recent years, deep networks have led to dramatic improvements in speech enhancement by framing it as a data-driven pattern recognition problem. In many modern enhancement systems, large amounts of data are used to train a deep network to estimate masks for complex-valued short-time Fourier transf…

Cited by 0SourceScholar
2019

The Phasebook: Building Complex Masks via Discrete Representations for Source Separation

ICASSP 2019accepted

Deep learning based speech enhancement and source separation systems have recently reached unprecedented levels of quality, to the point that performance is reaching a new ceiling. Most systems rely on estimating the magnitude of a target source, either directly or by computing a real-valued mask to…

Cited by 0SourceScholar
2018

An End-to-End Language-Tracking Speech Recognizer for Mixed-Language Speech

ICASSP 2018accepted

End-to-end automatic speech recognition (ASR) can significantly reduce the burden of developing ASR systems for new languages, by eliminating the need for linguistic information such as pronunciation dictionaries. This also creates an opportunity to build a monolithic multilingual ASR system with a…

Cited by 0SourceScholar
2018

End-to-End Multi-Speaker Speech Recognition

ICASSP 2018accepted

Current advances in deep learning have resulted in a convergence of methods across a wide range of tasks, opening the door for tighter integration of modules that were previously developed and optimized in isolation. Recent ground-breaking works have produced end-to-end deep network methods for both…

Cited by 0SourceScholar
2018

Multi-Channel Deep Clustering: Discriminative Spectral and Spatial Embeddings for Speaker-Independent Speech Separation

ICASSP 2018accepted

The recently-proposed deep clustering algorithm represents a fundamental advance towards solving the cocktail party problem in the single-channel case. When multiple microphones are available, spatial information can be leveraged to differentiate signals from different directions. This study combine…

Cited by 0SourceScholar
2018

Speaker Adaptation for Multichannel End-to-End Speech Recognition

ICASSP 2018accepted

Recent work on multichannel end-to-end automatic speech recognition (ASR) has shown that multichannel speech enhancement and speech recognition functions can be integrated into a deep neural network (DNN)-based system, and promising experimental results have been shown using the CHiME-4 and AMI corp…

Cited by 0SourceScholar
2017

Attention-Based Multimodal Fusion for Video Description

ICCV 2017poster

Current methods for video description are based on encoder-decoder sentence generation using recurrent neural networks (RNNs). Recent work has demonstrated the advantages of integrating temporal attention mechanisms into these models, in which the decoder network predicts each word in the descriptio…

Cited by 469PDFScholar
2017

Deep clustering and conventional networks for music separation: Stronger together

ICASSP 2017accepted

Deep clustering is the first method to handle general audio separation scenarios with multiple sources of the same type and an arbitrary number of sources, performing impressively in speaker-independent speech separation tasks. However, little is known about its effectiveness in other challenging si…

Cited by 0SourceScholar
2017

Deep long short-term memory adaptive beamforming networks for multichannel robust speech recognition

ICASSP 2017accepted

Far-field speech recognition in noisy and reverberant conditions remains a challenging problem despite recent deep learning breakthroughs. This problem is commonly addressed by acquiring a speech signal from multiple microphones and performing beamforming over them. In this paper, we propose to use…

Cited by 0SourceScholar
2017

Student-teacher network learning with enhanced features

ICASSP 2017accepted

Recent advances in distant-talking ASR research have confirmed that speech enhancement is an essential technique for improving the ASR performance, especially in the multichannel scenario. However, speech enhancement inevitably distorts speech signals, which can cause significant degradation when en…

Cited by 0SourceScholar
2016

Deep beamforming networks for multi-channel speech recognition

ICASSP 2016accepted

Despite the significant progress in speech recognition enabled by deep neural networks, poor performance persists in some scenarios. In this work, we focus on far-field speech recognition which remains challenging due to high levels of noise and reverberation in the captured speech signals. We propo…

Cited by 0SourceScholar
2016

Deep clustering: Discriminative embeddings for segmentation and separation

ICASSP 2016accepted

We address the problem of "cocktail-party" source separation in a deep learning framework called deep clustering. Previous deep network approaches to separation have shown promising performance in scenarios with a fixed number of sources, each belonging to a distinct signal class, such as speech and…

Cited by 0SourceScholar
2016

Minimum word error training of long short-term memory recurrent neural network language models for speech recognition

ICASSP 2016accepted

This paper describes minimum word error (MWE) training of recurrent neural network language models (RNNLMs) for speech recognition. RNNLMs are usually trained to minimize a cross entropy of estimated word probabilities against the correct word sequence, which corresponds to maximum likelihood criter…

Cited by 0SourceScholar
2015

Micbots: Collecting large realistic datasets for speech and audio research using mobile robots

ICASSP 2015accepted

Speech and audio signal processing research is a tale of data collection efforts and evaluation campaigns. Large benchmark datasets for automatic speech recognition (ASR) have been instrumental in the advancement of speech recognition technologies. However, when it comes to robust ASR, source separa…

Cited by 0SourceScholar
2015

Phase-sensitive and recognition-boosted speech separation using deep recurrent neural networks

ICASSP 2015accepted

Separation of speech embedded in non-stationary interference is a challenging problem that has recently seen dramatic improvements using deep network-based methods. Previous work has shown that estimating a masking function to be applied to the noisy spectrum is a viable approach that can be improve…

Cited by 0SourceScholar