← Search

Scott Wisdom

18 accepted papers

2025

Towards Sub-millisecond Latency Real-Time Speech Enhancement Models on Hearables

ICASSP 2025accepted

Low latency models are critical for real-time speech enhancement applications, such as hearing aids and hearables. However, the sub-millisecond latency space for resource-constrained hearables remains underexplored. We demonstrate speech enhancement using a computationally efficient minimum-phase FI…

Cited by 0SourceScholar
2022

Adapting Speech Separation to Real-World Meetings using Mixture Invariant Training

ICASSP 2022accepted

The recently-proposed mixture invariant training (MixIT) is an unsupervised method for training single-channel sound separation models because it does not require ground-truth isolated reference sources. In this paper, we investigate using MixIT to adapt a separation model on real far-field overlapp…

Cited by 26SourceScholar
2022

AudioScopeV2: Audio-Visual Attention Architectures for Calibrated Open-Domain On-Screen Sound Separation

ECCV 2022poster

"We introduce AudioScopeV2, a state-of-the-art universal audio-visual on-screen sound separation system which is capable of learning to separate sounds and associate them with on-screen objects by looking at in-the-wild videos. We identify several limitations of previous work on audio-visual on-scre…

2021

End-To-End Diarization for Variable Number of Speakers with Local-Global Networks and Discriminative Speaker Embeddings

ICASSP 2021accepted

We present an end-to-end deep network model that performs meeting diarization from single-channel audio recordings. End-to-end diarization models have the advantage of handling speaker overlap and enabling straightforward handling of discriminative training, unlike traditional clustering-based diari…

Cited by 0SourceScholar
2021

Into the Wild with AudioScope: Unsupervised Audio-Visual Separation of On-Screen Sounds

ICLR 2021poster

Recent progress in deep learning has enabled many advances in sound separation and visual scene understanding. However, extracting sound sources which are apparent in natural videos remains an open problem. In this work, we present AudioScope, a novel audio-visual sound separation framework that can…

Cited by 86SourcePDFScholar
2021

Sound Event Detection and Separation: A Benchmark on Desed Synthetic Soundscapes

ICASSP 2021accepted

We propose a benchmark of state-of-the-art sound event detection systems (SED). We design synthetic evaluation sets to focus on specific sound event detection challenges. We analyze the performance of the submissions to DCASE 2020 Task 4 as a function of time-related modifications (time position of…

Cited by 0SourceScholar
2021

What's all the Fuss about Free Universal Sound Separation Data?

ICASSP 2021accepted

We introduce the Free Universal Sound Separation (FUSS) dataset, a new corpus for experiments in separating mixtures of an unknown number of sounds from an open domain of sound types. The dataset consists of 23 hours of single-source audio data drawn from 357 classes, which are used to create mixtur…

Cited by 0SourceScholar
2020

Improving Universal Sound Separation Using Sound Classification

ICASSP 2020accepted

Deep learning approaches have recently achieved impressive performance on both audio source separation and sound classification. Most audio source separation approaches focus only on separating sources belonging to a restricted domain of source classes, such as speech and music. However, recent work…

Cited by 0SourceScholar
2020

Performance Study of a Convolutional Time-Domain Audio Separation Network for Real-Time Speech Denoising

ICASSP 2020accepted

Time-domain audio separation networks based on dilated temporal convolutions have recently been shown to perform very well compared to methods that are based on a time-frequency representation in speech separation tasks, even outperforming an oracle binary time-frequency mask of the speakers. This p…

Cited by 0SourceScholar
2020

Unsupervised Sound Separation Using Mixture Invariant Training

NeurIPS 2020spotlight

In recent years, rapid progress has been made on the problem of single-channel sound separation using supervised training of deep neural networks. In such supervised approaches, a model is trained to predict the component sources from synthetic mixtures created by adding up isolated ground-truth sou…

2019

Differentiable Consistency Constraints for Improved Deep Speech Enhancement

ICASSP 2019accepted

In recent years, deep networks have led to dramatic improvements in speech enhancement by framing it as a data-driven pattern recognition problem. In many modern enhancement systems, large amounts of data are used to train a deep network to estimate masks for complex-valued short-time Fourier transf…

Cited by 0SourceScholar
2017

Building recurrent networks by unfolding iterative thresholding for sequential sparse recovery

ICASSP 2017accepted

Historically, sparse methods and neural networks, particularly modern deep learning methods, have been relatively disparate areas. Sparse methods are typically used for signal enhancement, compression, and recovery, usually in an unsupervised framework, while neural networks commonly rely on a super…

Cited by 0SourceScholar
2016

Full-Capacity Unitary Recurrent Neural Networks

NeurIPS 2016poster

Recurrent neural networks are powerful models for processing sequential data, but they are generally plagued by vanishing and exploding gradient problems. Unitary recurrent neural networks (uRNNs), which use unitary recurrence matrices, have recently been proposed as a means to avoid these issues. H…