← Search

Nam Soo Kim

11 accepted papers

2025

Evidential-TTS: High Fidelity Zero-Shot Text-to-Speech Using Evidential Deep Learning

ICASSP 2025accepted

We propose Evidential-TTS, a novel zero-shot text-to-speech (TTS) system based on evidential deep learning (EDL). The model includes a length regulator to ensure precise alignment between phonemes and acoustic tokens. This module allows the evidential token generator to convert the aligned phoneme s…

Cited by 0SourceScholar
2025

FADEL: Uncertainty-aware Fake Audio Detection with Evidential Deep Learning

ICASSP 2025accepted

Recently, fake audio detection has gained significant attention, as advancements in speech synthesis and voice conversion have increased the vulnerability of automatic speaker verification (ASV) systems to spoofing attacks. A key challenge in this task is generalizing models to detect unseen, out-of…

Cited by 0SourceScholar
2023

EM-Network: Oracle Guided Self-distillation for Sequence Learning

ICML 2023poster

We introduce EM-Network, a novel self-distillation approach that effectively leverages target information for supervised sequence-to-sequence (seq2seq) learning. In contrast to conventional methods, it is trained with oracle guidance, which is derived from the target sequence. Since the oracle guida…

Cited by 3SourcePDFScholar
2023

Improving Learning Objectives for Speaker Verification from the Perspective of Score Comparison

ICASSP 2023accepted

Deep speaker embedding systems are usually trained with classification-based or end-to-end learning objectives. Popular end-to-end approaches utilize deep metric learning, which can be viewed as a few-shot classification objective. In this paper, we investigate the limit of conventional learning obj…

Cited by 0SourceScholar
2023

Multi-Resolution Sequence Aggregation and Model-Agnostic Framework for Time-Series Forecasting

ICASSP 2023accepted

In time-series forecasting, signals such as traffic volume collected in the real world are noisy and irregularly sampled due to sensor malfunctions, so it is difficult to make accurate prediction. To resolve such difficulty, downsampling can be used to reduce noise and allow capturing slow trend of…

Cited by 0SourceScholar
2020

Robust Front-End for Multi-Channel ASR using Flow-Based Density Estimation

IJCAI 2020poster

For multi-channel speech recognition, speech enhancement techniques such as denoising or dereverberation are conventionally applied as a front-end processor. Deep learning-based front-ends using such techniques require aligned clean and noisy speech pairs which are generally obtained via data simula…

Cited by 0SourcePDFScholar
2020

SoftFlow: Probabilistic Framework for Normalizing Flow on Manifolds

NeurIPS 2020poster

Flow-based generative models are composed of invertible transformations between two random variables of the same dimension. Therefore, flow-based models cannot be adequately trained if the dimension of the data distribution does not match that of the underlying target distribution. In this paper, we…

2017

Integrated DNN-based model adaptation technique for noise-robust speech recognition

ICASSP 2017accepted

Since the introduction of deep neural network (DNN)-based acoustic model, robust automatic speech recognition using DNN are being in research. Especially in model adaptation, the techniques utilizing auxiliary context features is known to be a promising technique. Recently, we proposed a technique w…

Cited by 0SourceScholar
2016

Two-stage noise aware training using asymmetric deep denoising autoencoder

ICASSP 2016accepted

Ever since the deep neural network (DNN)-based acoustic model appeared, the recognition performance of automatic speech recognition has been greatly improved. Due to this achievement, various researches on DNN-based technique for noise robustness are also in progress. Among these approaches, the noi…

Cited by 0SourceScholar