← Search

Seongkyu Mun

7 accepted papers

2025

AdaptVC: High Quality Voice Conversion with Adaptive Learning

ICASSP 2025accepted

The goal of voice conversion is to transform the speech of a source speaker to sound like that of a reference speaker while preserving the original content. A key challenge is to extract disentangled linguistic content from the source and voice style from the reference. While existing approaches lev…

Cited by 0SourceScholar
2024

Latent Filling: Latent Space Data Augmentation for Zero-Shot Speech Synthesis

ICASSP 2024accepted

Previous works in zero-shot text-to-speech (ZS-TTS) have attempted to enhance its systems by enlarging the training data through crowd-sourcing or augmenting existing speech data. However, the use of low-quality data has led to a decline in the overall system performance. To avoid such degradation,…

Cited by 0SourceScholar
2024

Mels-Tts : Multi-Emotion Multi-Lingual Multi-Speaker Text-To-Speech System Via Disentangled Style Tokens

ICASSP 2024accepted

This paper proposes a multi-emotion, multi-lingual, and multi-speaker text-to-speech (MELS-TTS) system, employing disentangled style tokens for effective emotion transfer. In speech encompassing various attributes, such as emotional state, speaker identity, and linguistic style, disentangling these…

Cited by 0SourceScholar
2021

Streaming End-to-End Speech Recognition with Jointly Trained Neural Feature Enhancement

ICASSP 2021accepted

In this paper, we present a streaming end-to-end speech recognition model based on Monotonic Chunkwise Attention (MoCha) jointly trained with enhancement layers. Even though the MoCha attention enables streaming speech recognition with recognition accuracy comparable to a full attention-based approa…

Cited by 0SourceScholar
2020

The Sound of My Voice: Speaker Representation Loss for Target Voice Separation

ICASSP 2020accepted

Content and style representations have been widely studied in the field of style transfer. In this paper, we propose a new loss function using speaker content representation for audio source separation, and we call it speaker representation loss. The objective is to extract the target speaker voice…

Cited by 0SourceScholar
2017

Deep Neural Network based learning and transferring mid-level audio features for acoustic scene classification

ICASSP 2017accepted

Deep Neural Network (DNN) based transfer learning has been shown to be effective in Visual Object Classification (VOC) for complementing the deficit of target domain training samples by adapting classifiers that have been pre-trained for other large-scaled DataBase (DB). Although there exists an abu…

Cited by 0SourceScholar