← Search

Soo-Whan Chung

8 accepted papers

2023

An Empirical Study on Speech Restoration Guided by Self-Supervised Speech Representation

ICASSP 2023accepted

Enhancing speech quality is an indispensable yet difficult task as it is often complicated by a range of degradation factors. In addition to additive noise, reverberation, clipping, and speech attenuation can all adversely affect speech quality. Speech restoration aims to recover speech components f…

Cited by 0SourceScholar
2023

Diffusion-Based Generative Speech Source Separation

ICASSP 2023accepted

We propose DiffSep, a new single channel source separation method based on score-matching of a stochastic differential equation (SDE). We craft a tailored continuous time diffusion-mixing process starting from the separated sources and converging to a Gaussian distribution centered on their mixture.…

Cited by 0SourceScholar
2022

Phase Continuity: Learning Derivatives of Phase Spectrum for Speech Enhancement

ICASSP 2022accepted

Modern neural speech enhancement models usually include various forms of phase information in their training loss terms, either explicitly or implicitly. However, these loss terms are typically designed to reduce the distortion of phase spectrum values at specific frequencies, which ensures they do…

Cited by 0SourceScholar
2021

Looking Into Your Speech: Learning Cross-Modal Affinity for Audio-Visual Speech Separation

CVPR 2021poster

In this paper, we address the problem of separating individual speech signals from videos using audio-visual neural processing. Most conventional approaches utilize frame-wise matching criteria to extract shared information between co-occurring audio and video. Thus, their performance heavily depend…

Cited by 56PDFScholar
2019

Gradient-based Active Learning Query Strategy for End-to-end Speech Recognition

ICASSP 2019accepted

In this paper, we propose an effective active learning query strategy for an automatic speech recognition system with the aim of reducing the training cost. Generally, training a deep neural network with supervised learning requires a massive amount of labeled data to obtain excellent performance. H…

Cited by 0SourceScholar
2019

Perfect Match: Improved Cross-modal Embeddings for Audio-visual Synchronisation

ICASSP 2019accepted

This paper proposes a new strategy for learning powerful cross-modal embeddings for audio-to-video synchronisation. Here, we set up the problem as one of cross-modal retrieval, where the objective is to find the most relevant audio segment given a short video clip. The method builds on the recent ad…

Cited by 0SourceScholar