← Search

Kyogu Lee

36 accepted papers

2025

Multidimensional Adaptive Coefficient for Inference Trajectory Optimization in Flow and Diffusion

ICML 2025poster

Flow and diffusion models have demonstrated strong performance and training stability across various tasks but lack two critical properties of simulation-based methods: freedom of dimensionality and adaptability to different inference trajectories. To address this limitation, we propose the Multidim…

Cited by 0SourcePDFScholar
2025

TokenSynth: A Token-based Neural Synthesizer for Instrument Cloning and Text-to-Instrument

ICASSP 2025accepted

Recent advancements in neural audio codecs have enabled the use of tokenized audio representations in various audio generation tasks, such as text-to-speech, text-to-audio, and text-to-music generation. Leveraging this approach, we propose TokenSynth, a novel neural synthesizer that utilizes a decod…

Cited by 0SourceScholar
2025

Variable Bitrate Residual Vector Quantization for Audio Coding

ICASSP 2025accepted

Recent state-of-the-art neural audio compression models have progressively adopted residual vector quantization (RVQ). Despite this success, these models employ a fixed number of codebooks per frame, which can be suboptimal in terms of rate-distortion tradeoff, particularly in scenarios with simple…

Cited by 12SourceScholar
2024

Differentiable Modal Synthesis for Physical Modeling of Planar String Sound and Motion Simulation

NeurIPS 2024poster

While significant advancements have been made in music generation and differentiable sound synthesis within machine learning and computer audition, the simulation of instrument vibration guided by physical laws has been underexplored. To address this gap, we introduce a novel model for simulating th…

Cited by 2SourcePDFScholar
2024

Emosical: An Emotion-Annotated Musical Theatre Dataset

EMNLP 2024finding

This paper presents Emosical, a multimodal open-source dataset of musical films. Emosical comprises video, vocal audio, text, and character identity paired samples with annotated emotion tags. Emosical provides rich emotion annotations for each sample by inferring the background story of the charact…

2024

Learning Semantic Information from Raw Audio Signal Using Both Contextual and Phonetic Representations

ICASSP 2024accepted

We propose a framework to learn semantics from raw audio signals using two types of representations, encoding contextual and phonetic information respectively. Specifically, we introduce a speech-to-unit processing pipeline that captures two types of representations with different time resolutions.…

Cited by 0SourceScholar
2023

Global HRTF Interpolation Via Learned Affine Transformation of Hyper-Conditioned Features

ICASSP 2023accepted

Estimating Head-Related Transfer Functions (HRTFs) of arbitrary source points is essential in immersive binaural audio rendering. Computing each individual’s HRTFs is challenging, as traditional approaches require expensive time and computational resources, while modern data-driven approaches are da…

Cited by 0SourceScholar
2023

Medleyvox: An Evaluation Dataset for Multiple Singing Voices Separation

ICASSP 2023accepted

Separation of multiple singing voices into each voice is a rarely studied area in music source separation research. The absence of a benchmark dataset has hindered its progress. In this paper, we present an evaluation dataset and provide baseline studies for multiple singing voices separation. First…

Cited by 12SourceScholar
2023

Music Mixing Style Transfer: A Contrastive Learning Approach to Disentangle Audio Effects

ICASSP 2023accepted

We propose an end-to-end music mixing style transfer system that converts the mixing style of an input multitrack to that of a reference song. This is achieved with an encoder pre-trained with a contrastive objective to extract only audio effects related information from a reference music recording.…

Cited by 0SourceScholar
2023

Show Me the Instruments: Musical Instrument Retrieval From Mixture Audio

ICASSP 2023accepted

As digital music production has become mainstream, the selection of appropriate virtual instruments plays a crucial role in determining the quality of music. To search the musical instrument samples or virtual instruments that make one’s desired sound, music producers use their ears to listen and co…

Cited by 0SourceScholar
2022

End-To-End Music Remastering System Using Self-Supervised And Adversarial Training

ICASSP 2022accepted

Mastering is an essential step in music production, but it is also a challenging task that has to go through the hands of experienced audio engineers, where they adjust tone, space, and volume of a song. Remastering follows the same technical process, in which the context lies in mastering a song fo…

Cited by 0SourceScholar
2021

Neural Analysis and Synthesis: Reconstructing Speech from Self-Supervised Representations

NeurIPS 2021poster

We present a neural analysis and synthesis (NANSY) framework that can manipulate the voice, pitch, and speed of an arbitrary speech signal. Most of the previous works have focused on using information bottleneck to disentangle analysis features for controllable synthesis, which usually results in p…

Cited by 178SourcePDFScholar
2021

Neural Audio Fingerprint for High-Specific Audio Retrieval Based on Contrastive Learning

ICASSP 2021accepted

Most of existing audio fingerprinting systems have limitations to be used for high-specific audio retrieval at scale. In this work, we generate a low-dimensional representation from a short unit segment of audio, and couple this fingerprint with a fast maximum inner-product search. To this end, we p…

Cited by 0SourceScholar
2021

Real-Time Denoising and Dereverberation wtih Tiny Recurrent U-Net

ICASSP 2021accepted

Modern deep learning-based models have seen outstanding performance improvement with speech enhancement tasks. The number of parameters of state-of-the-art models, however, is often too large to be deployed on devices for real-world applications. To this end, we propose Tiny Recurrent U-Net (TRU-Net…

Cited by 0SourceScholar
2021

Reverb Conversion Of Mixed Vocal Tracks Using An End-To-End Convolutional Deep Neural Network

ICASSP 2021accepted

Reverb plays a critical role in music production, where it provides listeners with spatial realization, timbre, and texture of the music. Yet, it is challenging to reproduce the musical reverb of a reference music track even by skilled engineers. In response, we propose an end-to-end system capable…

Cited by 0SourceScholar
2021

Room Adaptive Conditioning Method for Sound Event Classification in Reverberant Environments

ICASSP 2021accepted

Ensuring performance robustness for a variety of situations that can occur in real-world environments is one of the challenging tasks in sound event classification. One of the unpredictable and detrimental factors in performance, especially in indoor environments, is reverberation. To alleviate this…

Cited by 0SourceScholar
2020

Disentangling Timbre and Singing Style with Multi-Singer Singing Synthesis System

ICASSP 2020accepted

In this study, we define the identity of the singer with two independent concepts – timbre and singing style – and propose a multi-singer singing synthesis system that can model them separately. To this end, we extend our single-singer model into a multi-singer model in the following ways: first, we…

Cited by 0SourceScholar
2020

From Inference to Generation: End-to-end Fully Self-supervised Generation of Human Face from Speech

ICLR 2020poster

This work seeks the possibility of generating the human face from voice solely based on the audio-visual data without any human-labeled annotations. To this end, we propose a multi-modal learning framework that links the inference stage and generation stage. First, the inference networks are trained…

Cited by 34SourceScholar
2019

Enhancing Music Features by Knowledge Transfer from User-item Log Data

ICASSP 2019accepted

In this paper, we propose a novel method that exploits music listening log data for general-purpose music feature extraction. Despite the wealth of information available in the log data of user-item interactions, it has been mostly used for collaborative filtering to find similar items or users and…

Cited by 0SourceScholar
2019

Phase-Aware Speech Enhancement with Deep Complex U-Net

ICLR 2019poster

Most deep learning-based models for speech enhancement have mainly focused on estimating the magnitude of spectrogram while reusing the phase from noisy speech for reconstruction. This is due to the difficulty of estimating the phase of clean speech. To improve speech enhancement performance, we tac…

Cited by 476SourceScholar
2018

Cover Song Identification Using Song-to-Song Cross-Similarity Matrix with Convolutional Neural Network

ICASSP 2018accepted

In this paper, we propose a cover song identification algorithm using a convolutional neural network (CNN). We first train the CNN model to classify any non-/cover relationship, by feeding a cross-similarity matrix that is generated from a pair of songs as an input. Our main idea is to use the CNN o…

Cited by 0SourceScholar
2015

Informed source separation from monaural music with limited binary time-frequency annotation

ICASSP 2015accepted

This paper presents a novel informed audio source separation algorithm given a limited binary time-frequency annotation. Assuming that all the sources can be represented using a low-rank model, we derive an objective function to minimize the rank of the source spectrogram, and the error between the…

Cited by 0SourceScholar