← Search

Jaehun Kim

10 accepted papers

2026

Controllable Embedding Transformation for Mood-Guided Music Retrieval

ICASSP 2026poster

Music representations are the backbone of modern recommendation systems, powering playlist generation, similarity search, and personalized discovery. Yet most embeddings offer little control for adjusting a single musical attribute, e.g., changing only the mood of a track while preserving its genre…

Cited by 0SourcePDFScholar
2025

AdaptVC: High Quality Voice Conversion with Adaptive Learning

ICASSP 2025accepted

The goal of voice conversion is to transform the speech of a source speaker to sound like that of a reference speaker while preserving the original content. A key challenge is to extract disentangled linguistic content from the source and voice style from the reference. While existing approaches lev…

Cited by 0SourceScholar
2025

From Faces to Voices: Learning Hierarchical Representations for High-quality Video-to-Speech

CVPR 2025highlight

The objective of this study is to generate high-quality speech from silent talking face videos, a task also known as video-to-speech synthesis. A significant challenge in video-to-speech synthesis lies in the substantial modality gap between silent video and multi-faceted speech. In this paper, we…

2024

Fregrad: Lightweight and Fast Frequency-Aware Diffusion Vocoder

ICASSP 2024accepted

The goal of this paper is to generate realistic audio with a lightweight and fast diffusion-based vocoder, named FreGrad. Our framework consists of the following three key components: (1) We employ discrete wavelet transform that decomposes a complicated waveform into sub-band wavelets, which helps…

Cited by 0SourceScholar
2024

Let There Be Sound: Reconstructing High Quality Speech from Silent Videos

AAAI 2024technical

The goal of this work is to reconstruct high quality speech from lip motions alone, a task also known as lip-to-speech. A key challenge of lip-to-speech systems is the one-to-many mapping caused by (1) the existence of homophenes and (2) multiple speech variations, resulting in a mispronounced and o…

2024

On The Effect Of Data-Augmentation On Local Embedding Properties In The Contrastive Learning Of Music Audio Representations

ICASSP 2024accepted

Audio embeddings are crucial tools in understanding large catalogs of music. Typically embeddings are evaluated on the basis of the performance they provide in a wide range of downstream tasks, however few studies have investigated the local properties of the embedding spaces themselves which are im…

Cited by 0SourceScholar
2024

Seeing Through The Conversation: Audio-Visual Speech Separation Based on Diffusion Model

ICASSP 2024accepted

The objective of this work is to extract the target speaker’s voice from a mixture of voices using visual cues. Existing works on audio-visual speech separation have demonstrated their performance with promising intelligibility, but maintaining naturalness remains challenging. To address this issue,…

Cited by 0SourceScholar
2024

Similar but Faster: Manipulation of Tempo in Music Audio Embeddings for Tempo Prediction and Search

ICASSP 2024accepted

Audio embeddings enable large scale comparisons of the similarity of audio files for applications such as search and recommendation. Due to the subjectivity of audio similarity, it can be desirable to design systems that answer not only whether audio is similar, but similar in what way (e.g., wrt. t…

Cited by 0SourceScholar
2024

Tempo Estimation as Fully Self-Supervised Binary Classification

ICASSP 2024accepted

This paper addresses the problem of global tempo estimation in musical audio. Given that annotating tempo is time-consuming and requires certain musical expertise, few publicly available data sources exist to train machine learning models for this task. Towards alleviating this issue, we propose a f…

Cited by 0SourceScholar