← Search

Juhan Nam

25 accepted papers

2026

LLM2Fx-Tools: Tool Calling for Music Post-Production

ICLR 2026poster

This paper introduces LLM2Fx-Tools, a multimodal tool-calling framework that generates executable sequences of audio effects (Fx-chain) for music post-production. LLM2Fx-Tools uses a large language model (LLM) to understand audio inputs, select audio effects types, determine their order, and estimat…

Cited by 0SourcecodeScholar
2025

CLaMP 3: Universal Music Information Retrieval Across Unaligned Modalities and Unseen Languages

ACL 2025finding

CLaMP 3 is a unified framework developed to address challenges of cross-modal and cross-lingual generalization in music information retrieval. Using contrastive learning, it aligns all major music modalities–including sheet music, performance signals, and audio recordings–with multilingual text in a…

2025

Twenty-Five Years of MIR Research: Achievements, Practices, Evaluations, and Future Challenges

ICASSP 2025accepted

In this paper, we trace the evolution of Music Information Retrieval (MIR) over the past 25 years. While MIR gathers all kinds of research related to music informatics, a large part of it focuses on signal processing techniques for music data, fostering a close relationship with the IEEE Audio and A…

Cited by 0SourceScholar
2024

A Real-Time Lyrics Alignment System Using Chroma and Phonetic Features for Classical Vocal Performance

ICASSP 2024accepted

The goal of real-time lyrics alignment is to take live singing audio as input and to pinpoint the exact position within given lyrics on the fly. The task can benefit real-world applications such as the automatic subtitling of live concerts or operas. However, designing a real-time model poses a grea…

Cited by 0SourceScholar
2024

Enriching Music Descriptions with A Finetuned-LLM and Metadata for Text-to-Music Retrieval

ICASSP 2024accepted

Text-to-Music Retrieval, finding music based on a given natural language query, plays a pivotal role in content discovery within extensive music databases. To address this challenge, prior research has predominantly focused on a joint embedding of music audio and text, utilizing it to retrieve music…

Cited by 0SourceScholar
2024

Expressive Acoustic Guitar Sound Synthesis with an Instrument-Specific Input Representation and Diffusion Outpainting

ICASSP 2024accepted

Synthesizing performing guitar sound is a highly challenging task due to the polyphony and high variability in expression. Recently, deep generative models have shown promising results in synthesizing expressive polyphonic instrument sounds from music scores, often using a generic MIDI input. In thi…

Cited by 0SourceScholar
2024

K-pop Lyric Translation: Dataset, Analysis, and Neural-Modelling

COLING 2024main

Lyric translation, a field studied for over a century, is now attracting computational linguistics researchers. We identified two limitations in previous studies. Firstly, lyric translation studies have predominantly focused on Western genres and languages, with no previous study centering on K-pop…

2024

T-Foley: A Controllable Waveform-Domain Diffusion Model for Temporal-Event-Guided Foley Sound Synthesis

ICASSP 2024accepted

Foley sound, audio content inserted synchronously with videos, plays a critical role in the user experience of multimedia content. Recently, there has been active research in Foley sound synthesis, leveraging the advancements in deep generative models. However, such works mainly focus on replicating…

Cited by 0SourceScholar
2023

A Study of Audio Mixing Methods for Piano Transcription in Violin-Piano Ensembles

ICASSP 2023accepted

While piano music transcription models have shown high performance for solo piano recordings, their performance de-grades when applied to ensemble recordings. This study aims to analyze the impact of different data augmentation methods on piano transcription performance, specifically focusing on mix…

Cited by 0SourceScholar
2022

Pseudo-Label Transfer from Frame-Level to Note-Level in a Teacher-Student Framework for Singing Transcription from Polyphonic Music

ICASSP 2022accepted

Lack of large-scale note-level labeled data is the major obstacle to singing transcription from polyphonic music. We address the issue by using pseudo labels from vocal pitch estimation models given unlabeled data. The proposed method first converts the frame-level pseudo labels to note-level throug…

Cited by 20SourceScholar
2020

Disentangled Multidimensional Metric Learning for Music Similarity

ICASSP 2020accepted

Music similarity search is useful for a variety of creative tasks such as replacing one music recording with another recording with a similar "feel", a common task in video editing. For this task, it is typically necessary to define a similarity metric to compare one recording to another. Music simi…

Cited by 0SourceScholar
2020

Korean Singing Voice Synthesis Based on Auto-Regressive Boundary Equilibrium Gan

ICASSP 2020accepted

Singing voice synthesis is a generative task that involves not only multidimensional controls of a singer model such as phonetic modulation by lyrics and pitch control by music score but also expressive elements such as breath sounds and vibrato. Recently, end-to-end learning models based on generat…

Cited by 0SourceScholar
2019

Graph Neural Network for Music Score Data and Modeling Expressive Piano Performance

ICML 2019oral

Music score is often handled as one-dimensional sequential data. Unlike words in a text document, notes in music score can be played simultaneously by the polyphonic nature and each of them has its own duration. In this paper, we represent the unique form of musical score using graph neural network…

Cited by 78SourcePDFScholar