← Search

Yoonhyung Lee

4 accepted papers

2024

Boosting Speech Enhancement with Clean Self-Supervised Features Via Conditional Variational Autoencoders

ICASSP 2024accepted

Recently, Self-Supervised Features (SSF) trained on extensive speech datasets have shown significant performance gains across various speech processing tasks. Nevertheless, their effectiveness in Speech Enhancement (SE) systems is often suboptimal due to insufficient optimization for noisy environme…

Cited by 0SourceScholar
2022

Statistical inference with implicit SGD: proximal Robbins-Monro vs. Polyak-Ruppert

ICML 2022spotlight

The implicit stochastic gradient descent (ISGD), a proximal version of SGD, is gaining interest in the literature due to its stability over (explicit) SGD. In this paper, we conduct an in-depth analysis of the two modes of ISGD for smooth convex functions, namely proximal Robbins-Monro (proxRM) and…

Cited by 4SourcePDFScholar
2022

Varianceflow: High-Quality and Controllable Text-to-Speech using Variance Information via Normalizing Flow

ICASSP 2022accepted

There are two types of methods for non-autoregressive text-to-speech models to learn the one-to-many relationship between text and speech effectively. The first one is to use an advanced generative framework such as normalizing flow (NF). The second one is to use variance information such as pitch o…

Cited by 0SourceScholar
2021

Bidirectional Variational Inference for Non-Autoregressive Text-to-Speech

ICLR 2021poster

Although early text-to-speech (TTS) models such as Tacotron 2 have succeeded in generating human-like speech, their autoregressive architectures have several limitations: (1) They require a lot of time to generate a mel-spectrogram consisting of hundreds of steps. (2) The autoregressive speech gener…

Cited by 54SourcePDFScholar