← Search

Hoon-Young Cho

7 accepted papers

2025

Single-Channel Distance-Based Source Separation for Mobile GPU in Outdoor and Indoor Environments

ICASSP 2025accepted

This study emphasizes the significance of exploring distance-based source separation (DSS) in outdoor environments. Unlike existing studies that primarily focus on indoor settings, the proposed model is designed to capture the unique characteristics of outdoor audio sources. It incorporates advanced…

Cited by 0SourceScholar
2025

Text-Aware Adapter for Few-Shot Keyword Spotting

ICASSP 2025accepted

Recent advances in flexible keyword spotting (KWS) with text enrollment allow users to personalize keywords without uttering them during enrollment. However, there is still room for improvement in target keyword performance. In this work, we propose a novel few-shot transfer learning method, called…

Cited by 0SourceScholar
2024

FINALLY: fast and universal speech enhancement with studio-like quality

NeurIPS 2024poster

In this paper, we address the challenge of speech enhancement in real-world recordings, which often contain various forms of distortion, such as background noise, reverberation, and microphone artifacts. We revisit the use of Generative Adversarial Networks (GANs) for speech enhancement and theoreti…

2024

Latent Filling: Latent Space Data Augmentation for Zero-Shot Speech Synthesis

ICASSP 2024accepted

Previous works in zero-shot text-to-speech (ZS-TTS) have attempted to enhance its systems by enlarging the training data through crowd-sourcing or augmenting existing speech data. However, the use of low-quality data has led to a decline in the overall system performance. To avoid such degradation,…

Cited by 0SourceScholar
2024

Mels-Tts : Multi-Emotion Multi-Lingual Multi-Speaker Text-To-Speech System Via Disentangled Style Tokens

ICASSP 2024accepted

This paper proposes a multi-emotion, multi-lingual, and multi-speaker text-to-speech (MELS-TTS) system, employing disentangled style tokens for effective emotion transfer. In speech encompassing various attributes, such as emotional state, speaker identity, and linguistic style, disentangling these…

Cited by 0SourceScholar
2021

A Neural Text-to-Speech Model Utilizing Broadcast Data Mixed with Background Music

ICASSP 2021accepted

Recently, it has become easier to obtain speech data from various media such as the internet or YouTube, but directly utilizing them to train a neural text-to-speech (TTS) model is difficult. The proportion of clean speech is insufficient and the remainder includes background music. Even with the gl…

Cited by 0SourceScholar
2020

Detecting Mismatch Between Text Script and Voice-Over Using Utterance Verification Based on Phoneme Recognition Ranking

ICASSP 2020accepted

The purpose of this study is to detect the mismatch between text script and voice-over. For this, we present a novel utterance verification (UV) method, which calculates the degree of correspondence between a voice-over and the phoneme sequence of a script. We found that the phoneme recognition prob…

Cited by 0SourceScholar