← Search

Hong-Goo Kang

20 accepted papers

2025

LAMA-UT: Language Agnostic Multilingual ASR Through Orthography Unification and Language-Specific Transliteration

AAAI 2025technical

Building a universal multilingual automatic speech recognition (ASR) model that performs equitably across languages has long been a challenge due to its inherent difficulties. To address this task we introduce a Language-Agnostic Multilingual ASR pipeline through orthography Unification and language…

2025

StableQuant: Layer Adaptive Post-Training Quantization for Speech Foundation Models

ICASSP 2025accepted

In this paper, we propose StableQuant, a novel adaptive post-training quantization (PTQ) algorithm for widely used speech foundation models (SFMs). While PTQ has been successfully employed for compressing large language models (LLMs) due to its ability to bypass additional fine-tuning, directly appl…

Cited by 0SourceScholar
2023

End-to-End Neural Audio Coding in the MDCT Domain

ICASSP 2023accepted

Modern deep neural network (DNN)-based audio coding approaches utilize complicated non-linear functions (e.g., convolutional neural networks and non-linear activations), which leads to high complexity and memory usage. However, their decoded audio quality is still not much higher than that of signal…

Cited by 5SourceScholar
2023

Progressive Multi-Stage Neural Audio Codec with Psychoacoustic Loss and Discriminator

ICASSP 2023accepted

In this paper, we improve the efficiency of the progressive multi-stage neural audio codec (PR-Codec) by utilizing perceptually motivated training criteria. Although our baseline PR-Codec successfully reconstructs full-band signals by progressively decoding the pre-defined subband signals, transpare…

Cited by 6SourceScholar
2022

Adversarial Audio Synthesis Using a Harmonic-Percussive Discriminator

ICASSP 2022accepted

In this paper, we propose a discriminator design scheme for generative adversarial network-based audio signal generation. Unlike conventional discriminators that take an entire signal as input, our discriminator separates the audio signal into harmonic and percussive components and analyzes each com…

Cited by 0SourceScholar
2022

Phase Continuity: Learning Derivatives of Phase Spectrum for Speech Enhancement

ICASSP 2022accepted

Modern neural speech enhancement models usually include various forms of phase information in their training loss terms, either explicitly or implicitly. However, these loss terms are typically designed to reduce the distortion of phase spectrum values at specific frequencies, which ensures they do…

Cited by 0SourceScholar
2022

Progressive Multi-Stage Neural Audio Coding with Guided References

ICASSP 2022accepted

In this paper, we propose an effective multi-stage neural audio coding algorithm that encodes full-band audio signals (up to 20 kHz) using an end-to-end training criterion. By predefining several dyadic subband signals as training targets, we progressively encode input audio signals in each stage su…

Cited by 0SourceScholar
2021

Looking Into Your Speech: Learning Cross-Modal Affinity for Audio-Visual Speech Separation

CVPR 2021poster

In this paper, we address the problem of separating individual speech signals from videos using audio-visual neural processing. Most conventional approaches utilize frame-wise matching criteria to extract shared information between co-occurring audio and video. Thus, their performance heavily depend…

Cited by 56PDFScholar
2020

Emotional Speech Synthesis with Rich and Granularized Control

ICASSP 2020accepted

This paper proposes an effective emotion control method for an end-to-end text-to-speech (TTS) system. To flexibly control the distinct characteristic of a target emotion category, it is essential to determine embedding vectors representing the TTS input. We introduce an inter-to-intra emotional dis…

Cited by 0SourceScholar
2020

Improving LPCNET-Based Text-to-Speech with Linear Prediction-Structured Mixture Density Network

ICASSP 2020accepted

In this paper, we propose an improved LPCNet vocoder using a linear prediction (LP)-structured mixture density network (MDN). The recently proposed LPCNet vocoder has successfully achieved high-quality and lightweight speech synthesis systems by combining a vocal tract LP filter with a WaveRNN-based…

Cited by 0SourceScholar
2019

Gradient-based Active Learning Query Strategy for End-to-end Speech Recognition

ICASSP 2019accepted

In this paper, we propose an effective active learning query strategy for an automatic speech recognition system with the aim of reducing the training cost. Generally, training a deep neural network with supervised learning requires a massive amount of labeled data to obtain excellent performance. H…

Cited by 0SourceScholar
2019

Perfect Match: Improved Cross-modal Embeddings for Audio-visual Synchronisation

ICASSP 2019accepted

This paper proposes a new strategy for learning powerful cross-modal embeddings for audio-to-video synchronisation. Here, we set up the problem as one of cross-modal retrieval, where the objective is to find the most relevant audio segment given a short video clip. The method builds on the recent ad…

Cited by 0SourceScholar
2018

Dnn-Based Wireless Positioning in an Outdoor Environment

ICASSP 2018accepted

In this paper, we propose a deep learning based algorithm to estimate the position of an user by utilizing reference signal received power (RSRP) and the location of base stations. To obtain reliable results in a real communication environment, parameters were measured using commercially available b…

Cited by 0SourceScholar
2018

Modeling-By-Generation-Structured Noise Compensation Algorithm for Glottal Vocoding Speech Synthesis System

ICASSP 2018accepted

This paper proposes a novel noise compensation algorithm for a glottal excitation model in a deep learning (DL)-based speech synthesis system. To generate high-quality speech synthesis outputs, the balance between harmonic and noise components of the glottal excitation signal should be well-represen…

Cited by 0SourceScholar
2015

Improved time-frequency trajectory excitation modeling for a statistical parametric speech synthesis system

ICASSP 2015accepted

This paper proposes an improved time-frequency trajectory excitation (TFTE) modeling method for a statistical parametric speech synthesis system. The proposed approach overcomes the dimensional variation problem of the training process caused by the inherent nature of the pitch-dependent analysis pa…

Cited by 0SourceScholar