← Search

Hyeong-Seok Choi

8 accepted papers

2023

NANSY++: Unified Voice Synthesis with Neural Analysis and Synthesis

ICLR 2023poster

Various applications of voice synthesis have been developed independently despite the fact that they generate “voice” as output in common. In addition, most of the voice synthesis models still require a large number of audio data paired with annotated labels (e.g., text transcription and music score…

Cited by 62SourcePDFScholar
2023

Towards Trustworthy Phoneme Boundary Detection with Autoregressive Model and Improved Evaluation Metric

ICASSP 2023accepted

Phoneme boundary detection has been studied due to its central role in various speech applications. In this work, we point out that this task needs to be addressed not only by algorithmic way, but also by evaluation metric. To this end, we first propose a state-of-the-art phoneme boundary detector t…

Cited by 0SourceScholar
2021

Neural Analysis and Synthesis: Reconstructing Speech from Self-Supervised Representations

NeurIPS 2021poster

We present a neural analysis and synthesis (NANSY) framework that can manipulate the voice, pitch, and speed of an arbitrary speech signal. Most of the previous works have focused on using information bottleneck to disentangle analysis features for controllable synthesis, which usually results in p…

Cited by 178SourcePDFScholar
2021

Real-Time Denoising and Dereverberation wtih Tiny Recurrent U-Net

ICASSP 2021accepted

Modern deep learning-based models have seen outstanding performance improvement with speech enhancement tasks. The number of parameters of state-of-the-art models, however, is often too large to be deployed on devices for real-world applications. To this end, we propose Tiny Recurrent U-Net (TRU-Net…

Cited by 0SourceScholar
2021

Room Adaptive Conditioning Method for Sound Event Classification in Reverberant Environments

ICASSP 2021accepted

Ensuring performance robustness for a variety of situations that can occur in real-world environments is one of the challenging tasks in sound event classification. One of the unpredictable and detrimental factors in performance, especially in indoor environments, is reverberation. To alleviate this…

Cited by 0SourceScholar
2020

Disentangling Timbre and Singing Style with Multi-Singer Singing Synthesis System

ICASSP 2020accepted

In this study, we define the identity of the singer with two independent concepts – timbre and singing style – and propose a multi-singer singing synthesis system that can model them separately. To this end, we extend our single-singer model into a multi-singer model in the following ways: first, we…

Cited by 0SourceScholar
2020

From Inference to Generation: End-to-end Fully Self-supervised Generation of Human Face from Speech

ICLR 2020poster

This work seeks the possibility of generating the human face from voice solely based on the audio-visual data without any human-labeled annotations. To this end, we propose a multi-modal learning framework that links the inference stage and generation stage. First, the inference networks are trained…

Cited by 34SourceScholar
2019

Phase-Aware Speech Enhancement with Deep Complex U-Net

ICLR 2019poster

Most deep learning-based models for speech enhancement have mainly focused on estimating the magnitude of spectrogram while reusing the phase from noisy speech for reconstruction. This is due to the difficulty of estimating the phase of clean speech. To improve speech enhancement performance, we tac…

Cited by 476SourceScholar