← Search

Junghyun Koo

10 accepted papers

2026

Concept-TRAK: Understanding how diffusion models learn concepts through concept attribution

ICLR 2026poster

While diffusion models excel at image generation, their growing adoption raises critical concerns about copyright issues and model transparency. Existing attribution methods identify training examples influencing an entire image, but fall short in isolating contributions to specific elements, such a…

Cited by 0SourcecodeScholar
2026

LLM2Fx-Tools: Tool Calling for Music Post-Production

ICLR 2026poster

This paper introduces LLM2Fx-Tools, a multimodal tool-calling framework that generates executable sequences of audio effects (Fx-chain) for music post-production. LLM2Fx-Tools uses a large language model (LLM) to understand audio inputs, select audio effects types, determine their order, and estimat…

Cited by 0SourcecodeScholar
2025

Latent Diffusion Bridges for Unsupervised Musical Audio Timbre Transfer

ICASSP 2025accepted

Music timbre transfer is a challenging task that involves modifying the timbral characteristics of an audio signal while preserving its melodic structure. In this paper, we propose a novel method based on dual diffusion bridges, trained using the CocoChorales Dataset, which consists of unpaired mono…

Cited by 0SourceScholar
2025

TokenSynth: A Token-based Neural Synthesizer for Instrument Cloning and Text-to-Instrument

ICASSP 2025accepted

Recent advancements in neural audio codecs have enabled the use of tokenized audio representations in various audio generation tasks, such as text-to-speech, text-to-audio, and text-to-music generation. Leveraging this approach, we propose TokenSynth, a novel neural synthesizer that utilizes a decod…

Cited by 0SourceScholar
2025

Variable Bitrate Residual Vector Quantization for Audio Coding

ICASSP 2025accepted

Recent state-of-the-art neural audio compression models have progressively adopted residual vector quantization (RVQ). Despite this success, these models employ a fixed number of codebooks per frame, which can be suboptimal in terms of rate-distortion tradeoff, particularly in scenarios with simple…

Cited by 12SourceScholar
2023

Music Mixing Style Transfer: A Contrastive Learning Approach to Disentangle Audio Effects

ICASSP 2023accepted

We propose an end-to-end music mixing style transfer system that converts the mixing style of an input multitrack to that of a reference song. This is achieved with an encoder pre-trained with a contrastive objective to extract only audio effects related information from a reference music recording.…

Cited by 0SourceScholar
2022

End-To-End Music Remastering System Using Self-Supervised And Adversarial Training

ICASSP 2022accepted

Mastering is an essential step in music production, but it is also a challenging task that has to go through the hands of experienced audio engineers, where they adjust tone, space, and volume of a song. Remastering follows the same technical process, in which the context lies in mastering a song fo…

Cited by 0SourceScholar
2021

Reverb Conversion Of Mixed Vocal Tracks Using An End-To-End Convolutional Deep Neural Network

ICASSP 2021accepted

Reverb plays a critical role in music production, where it provides listeners with spatial realization, timbre, and texture of the music. Yet, it is challenging to reproduce the musical reverb of a reference music track even by skilled engineers. In response, we propose an end-to-end system capable…

Cited by 0SourceScholar
2020

Disentangling Timbre and Singing Style with Multi-Singer Singing Synthesis System

ICASSP 2020accepted

In this study, we define the identity of the singer with two independent concepts – timbre and singing style – and propose a multi-singer singing synthesis system that can model them separately. To this end, we extend our single-singer model into a multi-singer model in the following ways: first, we…

Cited by 0SourceScholar