← Search

Inseon Jang

9 accepted papers

2023

End-to-End Neural Audio Coding in the MDCT Domain

ICASSP 2023accepted

Modern deep neural network (DNN)-based audio coding approaches utilize complicated non-linear functions (e.g., convolutional neural networks and non-linear activations), which leads to high complexity and memory usage. However, their decoded audio quality is still not much higher than that of signal…

Cited by 0SourceScholar
2023

Individual Sub-Band Estimation Approach to Bandwidth Extension and Enhancement of Coded Speech

ICASSP 2023accepted

The streaming Sound EnhAncement Network (SEANet) has demonstrated impressive performance for speech bandwidth extension (BWE) with low latency and computational complexity. Although the streaming SEANet was designed for voice communication systems, it was not tested with decoded signals that include…

Cited by 0SourceScholar
2023

Native Multi-Band Audio Coding Within Hyper-Autoencoded Reconstruction Propagation Networks

ICASSP 2023accepted

Spectral sub-bands do not portray the same perceptual relevance. In audio coding, it is therefore desirable to have independent control over each of the constituent bands so that bitrate assignment and signal reconstruction can be achieved efficiently. In this work, we present a novel neural audio c…

Cited by 0SourceScholar
2023

Progressive Multi-Stage Neural Audio Codec with Psychoacoustic Loss and Discriminator

ICASSP 2023accepted

In this paper, we improve the efficiency of the progressive multi-stage neural audio codec (PR-Codec) by utilizing perceptually motivated training criteria. Although our baseline PR-Codec successfully reconstructs full-band signals by progressively decoding the pre-defined subband signals, transpare…

Cited by 0SourceScholar
2022

Adversarial Audio Synthesis Using a Harmonic-Percussive Discriminator

ICASSP 2022accepted

In this paper, we propose a discriminator design scheme for generative adversarial network-based audio signal generation. Unlike conventional discriminators that take an entire signal as input, our discriminator separates the audio signal into harmonic and percussive components and analyzes each com…

Cited by 0SourceScholar
2022

Progressive Multi-Stage Neural Audio Coding with Guided References

ICASSP 2022accepted

In this paper, we propose an effective multi-stage neural audio coding algorithm that encodes full-band audio signals (up to 20 kHz) using an end-to-end training criterion. By predefining several dyadic subband signals as training targets, we progressively encode input audio signals in each stage su…

Cited by 0SourceScholar
2020

Emotional Speech Synthesis with Rich and Granularized Control

ICASSP 2020accepted

This paper proposes an effective emotion control method for an end-to-end text-to-speech (TTS) system. To flexibly control the distinct characteristic of a target emotion category, it is essential to determine embedding vectors representing the TTS input. We introduce an inter-to-intra emotional dis…

Cited by 0SourceScholar