← Search

Hyeongju Kim

8 accepted papers

2026

Robust TTS Training via Self-Purifying Flow Matching for the WildSpoof 2026 TTS Track

ICASSP 2026poster

This paper presents a lightweight text-to-speech (TTS) system developed for the WildSpoof Challenge TTS Track. Our approach fine-tunes the recently released open-weight TTS model, \textit{Supertonic}\footnote{\url{https://github.com/supertone-inc/supertonic}}, with Self-Purifying Flow Matching (SPFM…

Cited by 0SourcePDFScholar
2023

NANSY++: Unified Voice Synthesis with Neural Analysis and Synthesis

ICLR 2023poster

Various applications of voice synthesis have been developed independently despite the fact that they generate “voice” as output in common. In addition, most of the voice synthesis models still require a large number of audio data paired with annotated labels (e.g., text transcription and music score…

Cited by 62SourcePDFScholar
2023

Towards Trustworthy Phoneme Boundary Detection with Autoregressive Model and Improved Evaluation Metric

ICASSP 2023accepted

Phoneme boundary detection has been studied due to its central role in various speech applications. In this work, we point out that this task needs to be addressed not only by algorithmic way, but also by evaluation metric. To this end, we first propose a state-of-the-art phoneme boundary detector t…

Cited by 0SourceScholar
2020

Robust Front-End for Multi-Channel ASR using Flow-Based Density Estimation

IJCAI 2020poster

For multi-channel speech recognition, speech enhancement techniques such as denoising or dereverberation are conventionally applied as a front-end processor. Deep learning-based front-ends using such techniques require aligned clean and noisy speech pairs which are generally obtained via data simula…

Cited by 0SourcePDFScholar
2020

SoftFlow: Probabilistic Framework for Normalizing Flow on Manifolds

NeurIPS 2020poster

Flow-based generative models are composed of invertible transformations between two random variables of the same dimension. Therefore, flow-based models cannot be adequately trained if the dimension of the data distribution does not match that of the underlying target distribution. In this paper, we…