← Search

Gordon Wichern

26 accepted papers

2025

30+ Years of Source Separation Research: Achievements and Future Challenges

ICASSP 2025accepted

Source separation (SS) of acoustic signals is a research field that emerged in the mid-1990s and has flourished ever since. On the occasion of ICASSP’s 50<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">th</sup> anniversary, we review the major contribut…

Cited by 23SourceScholar
2025

Keeping the Balance: Anomaly Score Calculation for Domain Generalization

ICASSP 2025accepted

Emitted sounds may drastically change when using different microphones, when properties of the sound sources change, or when recording in different acoustic environments. Ideally, anomalous sound detection (ASD) systems should be able to generalize well to unseen target domains by only providing a f…

Cited by 0SourceScholar
2025

Leveraging Audio-Only Data for Text-Queried Target Sound Extraction

ICASSP 2025accepted

The goal of text-queried target sound extraction (TSE) is to extract from a mixture a sound source specified with a natural-language caption. While it is preferable to have access to large-scale text-audio pairs to address a variety of text queries, the limited number of available high-quality text-…

Cited by 0SourceScholar
2025

No Class Left Behind: A Closer Look at Class Balancing for Audio Tagging

ICASSP 2025accepted

Large-scale audio tagging datasets like AudioSet usually suffer from severe class imbalance comprising many audio examples for common sound classes but only few examples of rare sound classes. The latter, however, may yet be equally or even more important to recognize. Therefore, it is common practi…

Cited by 0SourceScholar
2025

Retrieval-Augmented Neural Field for HRTF Upsampling and Personalization

ICASSP 2025accepted

Head-related transfer functions (HRTFs) with dense spatial grids are desired for immersive binaural audio generation, but their recording is time-consuming. Although HRTF spatial upsampling has shown remarkable progress with neural fields, spatial upsampling only from a few measured directions, e.g.…

Cited by 0SourceScholar
2025

Task-Aware Unified Source Separation

ICASSP 2025accepted

Several attempts have been made to handle multiple source separation tasks such as speech enhancement, speech separation, sound event separation, music source separation (MSS), or cinematic audio source separation (CASS) with a single model. These models are trained on large-scale data including spe…

Cited by 0SourceScholar
2024

Generation or Replication: Auscultating Audio Latent Diffusion Models

ICASSP 2024accepted

The introduction of audio latent diffusion models possessing the ability to generate realistic sound clips on demand from a text description has the potential to revolutionize how we work with audio. In this work, we make an initial attempt at understanding the inner workings of audio latent diffusi…

Cited by 0SourceScholar
2024

Improving Audio Captioning Models with Fine-Grained Audio Features, Text Embedding Supervision, and LLM Mix-Up Augmentation

ICASSP 2024accepted

Automated audio captioning (AAC) aims to generate informative descriptions for various sounds from nature and/or human activities. In recent years, AAC has quickly attracted research interest, with state-of-the-art systems now relying on a sequence-to-sequence (seq2seq) backbone powered by strong mo…

Cited by 0SourceScholar
2024

NIIRF: Neural IIR Filter Field for HRTF Upsampling and Personalization

ICASSP 2024accepted

Head-related transfer functions (HRTFs) are important for immersive audio, and their spatial interpolation has been studied to upsample finite measurements. Recently, neural fields (NFs) which map from sound source direction to HRTF have gained attention. Existing NF-based methods focused on estimat…

Cited by 0SourceScholar
2024

NeuroHeed+: Improving Neuro-Steered Speaker Extraction with Joint Auditory Attention Detection

ICASSP 2024accepted

Neuro-steered speaker extraction aims to extract the listener’s brainattended speech signal from a multi-talker speech signal, in which the attention is derived from the cortical activity. This activity is usually recorded using electroencephalography (EEG) devices. Though promising, current methods…

Cited by 0SourceScholar
2023

Latent Iterative Refinement for Modular Source Separation

ICASSP 2023accepted

Traditional source separation approaches train deep neural network models end-to-end with all the data available at once by minimizing the empirical risk on the whole training set. On the inference side, after training the model, the user fetches a static computation graph and runs the full model on…

Cited by 0SourceScholar
2023

Optimal Condition Training for Target Source Separation

ICASSP 2023accepted

Recent research has shown remarkable performance in leveraging multiple extraneous conditional and non-mutually-exclusive semantic concepts for sound source separation, allowing the flexibility to extract a given target source based on multiple different queries. In this work, we propose a new optim…

Cited by 0SourceScholar
2023

Reverberation as Supervision For Speech Separation

ICASSP 2023accepted

This paper proposes reverberation as supervision (RAS), a novel unsupervised loss function for single-channel reverberant speech separation. Prior methods for unsupervised separation required the synthesis of mixtures of mixtures or assumed the existence of a teacher model, making them difficult to…

Cited by 9SourceScholar
2022

Locate This, Not that: Class-Conditioned Sound Event DOA Estimation

ICASSP 2022accepted

Existing systems for sound event localization and detection (SELD) typically operate by estimating a source location for all classes at every time instant. In this paper, we propose an alternative class-conditioned SELD model for situations where we may not be interested in localizing all classes al…

Cited by 0SourceScholar
2022

The Cocktail Fork Problem: Three-Stem Audio Separation for Real-World Soundtracks

ICASSP 2022accepted

The cocktail party problem aims at isolating any source of interest within a complex acoustic scene, and has long inspired audio source separation research. Recent efforts have mainly focused on separating speech from noise, speech from speech, musical instruments from each other, or sound events fr…

Cited by 0SourceScholar
2021

Transcription Is All You Need: Learning To Separate Musical Mixtures With Score As Supervision

ICASSP 2021accepted

Most music source separation systems require large collections of isolated sources for training, which can be difficult to obtain. In this work, we use musical scores, which are comparatively easy to obtain, as a weak label for training a source separation system. In contrast with previous score-inf…

Cited by 13SourceScholar
2020

WHAMR!: Noisy and Reverberant Single-Channel Speech Separation

ICASSP 2020accepted

While significant advances have been made with respect to the separation of overlapping speech signals, studies have been largely constrained to mixtures of clean, near anechoic speech, not representative of many real-world scenarios. Although the WHAM! dataset introduced noise to the ubiquitous wsj…

Cited by 0SourceScholar
2019

Bootstrapping Single-channel Source Separation via Unsupervised Spatial Clustering on Stereo Mixtures

ICASSP 2019accepted

Separating an audio scene into isolated sources is a fundamental problem in computer audition, analogous to image segmentation in visual scene analysis. Source separation systems based on deep learning are currently the most successful approaches for solving the underdetermined separation problem, w…

Cited by 0SourceScholar
2019

Class-conditional Embeddings for Music Source Separation

ICASSP 2019accepted

Isolating individual instruments in a musical mixture has a myriad of potential applications, and seems imminently achievable given the levels of performance reached by recent deep learning methods. While most musical source separation techniques learn an independent model for each instrument, we pr…

Cited by 48SourceScholar
2019

End-to-end Audio Visual Scene-aware Dialog Using Multimodal Attention-based Video Features

ICASSP 2019accepted

In order for machines interacting with the real world to have conversations with users about the objects and events around them, they need to understand dynamic audiovisual scenes. The recent revolution of neural network models allows us to combine various modules into a single end-to-end differenti…

Cited by 0SourceScholar
2019

Teacher-student Deep Clustering for Low-delay Single Channel Speech Separation

ICASSP 2019accepted

The recently-proposed deep clustering algorithm introduced significant advances in monaural speaker-independent multi-speaker speech separation. Deep clustering operates on magnitude spectro-grams using bidirectional recurrent networks and K-means clustering, both of which require offline operation,…

Cited by 0SourceScholar
2019

The Phasebook: Building Complex Masks via Discrete Representations for Source Separation

ICASSP 2019accepted

Deep learning based speech enhancement and source separation systems have recently reached unprecedented levels of quality, to the point that performance is reaching a new ceiling. Most systems rely on estimating the magnitude of a target source, either directly or by computing a real-valued mask to…

Cited by 0SourceScholar