← Search

Minje Kim

36 accepted papers

2026

GENCHO: ROOM IMPULSE RESPONSE GENERATION FROM REVERBERANT SPEECH AND TEXT VIA DIFFUSION TRANSFORMERS

ICASSP 2026oral

Blind room impulse response (RIR) estimation is a core task for capturing and transferring acoustic properties; yet existing methods often suffer from limited modeling capability and degraded performance under unseen conditions. Moreover, emerging generative audio applications call for more flexible…

Cited by 0SourcePDFScholar
2026

PROMPTSEP: GENERATIVE AUDIO SEPARATION VIA MULTIMODAL PROMPTING

ICASSP 2026oral

Recent breakthroughs in language-queried audio source separation (LASS) have shown that generative models can achieve higher separation audio quality than traditional masking-based approaches. However, two key limitations restrict their practical use: (1) users often require operations beyond separa…

Cited by 0SourcePDFScholar
2026

Privacy-Preserving Argumentative Explanations (Student Abstract)

AAAI 2026technical

We propose a framework for privacy-preserving argumentative explanations using homomorphic encryption. This method applies the Cheon-Kim-Kim-Song scheme, along with a soft k-means adapted for encrypted computation, to generate explanations without exposing sensitive data. By leveraging GPU accelerat

Cited by 0SourcePDFScholar
2025

Perceptual Audio Coding: A 40-Year Historical Perspective

ICASSP 2025accepted

In the history of audio and acoustic signal processing, perceptual audio coding has certainly excelled as a bright success story by its ubiquitous deployment in virtually all digital media devices, such as computers, tablets, mobile phones, set-top-boxes, and digital radios. From a technology perspe…

Cited by 0SourceScholar
2025

SRHand: Super-Resolving Hand Images and 3D Shapes via View/Pose-aware Neural Image Representations and Explicit Meshes

NeurIPS 2025poster

Reconstructing detailed hand avatars plays a crucial role in various applications. While prior works have focused on capturing high-fidelity hand geometry, they heavily rely on high-resolution multi-view image inputs and struggle to generalize on low-resolution images. Multi-view image super-resolut…

Cited by 0SourceScholar
2024

A Benchmark Dataset for Collaborative SLAM in Service Environments

RA-L 2024

We introduce a new multi-modal collaborative SLAM (C-SLAM) dataset for multiple service robots in various indoor service environments, called <monospace xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">C</monospace>-SLAM dataset in <monospace xmlns:mml="http:

Cited by 4SourcecodeScholar
2024

BiTT: Bi-directional Texture Reconstruction of Interacting Two Hands from a Single Image

CVPR 2024poster

Creating personalized hand avatars is important to offer a realistic experience to users on AR / VR platforms. While most prior studies focused on reconstructing 3D hand shapes some recent work has tackled the reconstruction of hand textures on top of shapes. However these methods are often limited…

2024

Dense Hand-Object(HO) GraspNet with Full Grasping Taxonomy and Dynamics

ECCV 2024poster

"Existing datasets for 3D hand-object interaction are limited either in the data cardinality, data variations in interaction scenarios, or the quality of annotations. In this work, we present a comprehensive new training dataset for hand-object interaction called HOGraspNet. It is the only real data…

2024

Scalable and Efficient Speech Enhancement Using Modified Cold Diffusion: A Residual Learning Approach

ICASSP 2024accepted

We introduce flexibility to the supervised learning-based speech enhancement framework to achieve scalable and efficient speech enhancement (SESE). To this end, SESE conducts a series of segmented speech enhancement inference routines, each of which incrementally improves the result of its preceding…

Cited by 0SourceScholar
2023

Native Multi-Band Audio Coding Within Hyper-Autoencoded Reconstruction Propagation Networks

ICASSP 2023accepted

Spectral sub-bands do not portray the same perceptual relevance. In audio coding, it is therefore desirable to have independent control over each of the constituent bands so that bitrate assignment and signal reconstruction can be achieved efficiently. In this work, we present a novel neural audio c…

Cited by 0SourceScholar
2023

Neural Feature Predictor and Discriminative Residual Coding for Low-Bitrate Speech Coding

ICASSP 2023accepted

Low and ultra-low-bitrate neural speech codecs achieved unprecedented coding gain by generating speech signals from compact features. This paper introduces additional coding efficiency in speech coding by reducing the temporal redundancy existing in the frame-level feature sequence via a feature pre…

Cited by 0SourceScholar
2023

The Potential of Neural Speech Synthesis-Based Data Augmentation for Personalized Speech Enhancement

ICASSP 2023accepted

With the advances in deep learning, speech enhancement systems benefited from large neural network architectures and achieved state-of-the-art quality. However, speaker-agnostic methods are not always desirable, both in terms of quality and their complexity, when they are to be used in a resource-co…

Cited by 0SourceScholar
2022

Bloom-Net: Blockwise Optimization for Masking Networks Toward Scalable and Efficient Speech Enhancement

ICASSP 2022accepted

In this paper, we present a blockwise optimization method for masking-based networks (BLOOM-Net) for training scalable speech enhancement networks. Here, we design our network with a residual learning scheme and train the internal separator blocks sequentially to obtain a scalable masking-based deep…

Cited by 0SourceScholar
2022

Deep Adaptive Aec: Hybrid of Deep Learning and Adaptive Acoustic Echo Cancellation

ICASSP 2022accepted

In this paper we integrate classic adaptive filtering algorithms with modern deep learning to propose a new approach called deep adaptive AEC. The main idea is to represent the linear adaptive algorithm as a differentiable layer within a deep neural network (DNN) framework. This enables the gradient…

Cited by 0SourceScholar
2022

Don't Separate, Learn To Remix: End-To-End Neural Remixing With Joint Optimization

ICASSP 2022accepted

The task of manipulating the level and/or effects of individual instruments to recompose a mixture of recordings, or remixing, is common across a variety of applications such as music production, audio-visual post-production, podcasts, and more. This process, however, traditionally requires access t…

Cited by 0SourceScholar
2022

Upmixing Via Style Transfer: A Variational Autoencoder for Disentangling Spatial Images And Musical Content

ICASSP 2022accepted

In the stereo-to-multichannel upmixing problem for music, one of the main tasks is to set the directionality of the instrument sources in the multichannel rendering results. In this paper, we propose a modified variational autoencoder model that learns a latent space to describe the spatial images i…

Cited by 0SourceScholar
2020

A Dual-Staged Context Aggregation Method towards Efficient End-to-End Speech Enhancement

ICASSP 2020accepted

In speech enhancement, an end-to-end deep neural network converts a noisy speech signal to a clean speech directly in the time domain without time-frequency transformation or mask estimation. However, aggregating contextual information from a high-resolution time domain signal with an affordable mod…

Cited by 0SourceScholar
2020

Boosted Locality Sensitive Hashing: Discriminative Binary Codes for Source Separation

ICASSP 2020accepted

Speech enhancement tasks have seen significant improvements with the advance of deep learning technology, but with the cost of increased computational complexity. In this study, we propose an adaptive boosting approach to learning locality sensitive hash codes, which represent audio spectra efficien…

Cited by 0SourceScholar
2020

Deep Autotuner: A Pitch Correcting Network for Singing Performances

ICASSP 2020accepted

We introduce a data-driven approach to automatic pitch correction of solo singing performances. The proposed approach predicts note-wise pitch shifts from the relationship between the respective spectrograms of the singing and accompaniment. This approach differs from commercial systems, where vocal…

Cited by 0SourceScholar
2020

Efficient and Scalable Neural Residual Waveform Coding with Collaborative Quantization

ICASSP 2020accepted

Scalability and efficiency are desired in neural speech codecs, which supports a wide range of bitrates for applications on various devices. We propose a collaborative quantization (CQ) scheme to jointly learn the codebook of LPC coefficients and the corresponding residuals. CQ does not simply shoeh…

Cited by 0SourceScholar
2019

Incremental Binarization on Recurrent Neural Networks for Single-channel Source Separation

ICASSP 2019accepted

This paper proposes a Bitwise Gated Recurrent Unit (BGRU) network for the single-channel source separation task. Recurrent Neural Networks (RNN) require several sets of weights within its cells, which significantly increases the computational cost compared to the fully-connected networks. To mitigat…

Cited by 0SourceScholar
2019

Intonation: A Dataset of Quality Vocal Performances Refined by Spectral Clustering on Pitch Congruence

ICASSP 2019accepted

We introduce the "Intonation" dataset of amateur vocal performances with a tendency for good intonation, collected from Smule, Inc. The dataset can be used for music information retrieval tasks such as autotuning, query by humming, and singing style analysis. It is available upon request on the Stan…

Cited by 0SourceScholar
2018

Bitwise Source Separation on Hashed Spectra: An Efficient Posterior Estimation Scheme Using Partial Rank Order Metrics

ICASSP 2018accepted

This paper proposes an efficient bitwise solution to the single-channel source separation task. Most dictionary-based source separation algorithms rely on iterative update rules during the run time, which becomes computationally costly especially when we employ an overcomplete dictionary and sparse…

Cited by 0SourceScholar
2017

Towards expressive instrument synthesis through smooth frame-by-frame reconstruction: From string to woodwind

ICASSP 2017accepted

We consider the task of mapping the performance of a musical excerpt on one instrument to another. Our focus is on excitation-continuous instruments, where pitch, amplitude, spectrum, and time envelope are controlled continuously by the player. The synthesized instrument should follow the target ins…

Cited by 0SourceScholar
2016

Efficient neighborhood-based topic modeling for collaborative audio enhancement on massive crowdsourced recordings

ICASSP 2016accepted

Collaborative Audio Enhancement (CAE) aims at separating a dominant source from crowdsourced recordings of a scene. This paper proposes a CAE setup as a big ad-hoc microphone array problem, assuming hundreds of sensors scattered over a large scene, e.g. a concert hall or a street riot. An important…

Cited by 0SourceScholar
2015

Efficient manifold preserving audio source separation using locality sensitive hashing

ICASSP 2015accepted

We propose an efficient technique to learn probabilistic hierarchical topic models that are designed to preserve the manifold structure of audio data. The consideration of the data manifold is important, as it has been shown to provide superior performance in certain audio applications such as sourc…

Cited by 0SourceScholar