← Search

Chanwoo Kim

28 accepted papers

2026

LEGO: Latent-Space Exploration for Geometry-Aware Optimization of Humanoid Kinematic Design

ICRA 2026poster

Designing robot morphologies and kinematics has traditionally relied on human intuition, with little systematic foundation. Motion–design co-optimization offers a promising path toward automation, but two major challenges remain: (i) the vast, unstructured design space and (ii) the difficulty of con…

2026

Learning Social Navigation from Positive and Negative Demonstrations and Rule-Based Specifications

ICRA 2026poster

Mobile robot navigation in dynamic human environments requires policies that balance adaptability to diverse behaviors with compliance to safety constraints. We hypothesize that integrating data-driven rewards with rule-based objectives enables navigation policies to achieve a more effective balance…

2026

SurrogateSHAP: Training-Free Contributor Attribution for Text-to-Image (T2I) Models

ICML 2026poster

As Text-to-Image (T2I) diffusion models are increasingly used in real-world creative workflows, a principled framework for valuing contributors who provide a collection of data is essential for fair compensation and sustainable data marketplaces. While the Shapley value offers a theoretically ground…

Cited by 0SourceScholar
2025

Acceleration Measurement-Free Dissipative Disturbance Observer for Robotic Manipulators

RA-L 2025

In this letter, we propose an Acceleration Measurement-Free Dissipative Disturbance Observer (AFDDO) for robotic manipulators, designed to estimate external disturbances robustly without requiring angular acceleration measurements and matrix inversion. By leveraging dissipativity theory, the AFDDO a

Cited by 1SourceScholar
2025

Adaptive Hybrid Control for Backlash-Like Hysteresis and Marker-Based Pose Estimation in Endoscopic Robots

RA-L 2025

Nonlinear backlash-like hysteresis in tendon-sheath mechanism based endoscopic surgical robots introduces significant motion errors, limiting precision in surgical tasks. Existing methods struggle to compensate for these errors in real time, particularly without relying on distal-end sensors. To add

Cited by 0SourceScholar
2025

An Efficient Framework for Crediting Data Contributors of Diffusion Models

ICLR 2025poster

As diffusion models are deployed in real-world settings and their performance driven by training data, appraising the contribution of data contributors is crucial to creating incentives for sharing quality data and to implementing policies for data compensation. Depending on the use case, model perf…

Cited by 0SourcePDFScholar
2025

CellCLIP - Learning Perturbation Effects in Cell Painting via Text-Guided Contrastive Learning

NeurIPS 2025poster

High-content screening (HCS) assays based on high-throughput microscopy techniques such as Cell Painting have enabled the interrogation of cells' morphological responses to perturbations at an unprecedented scale. The collection of such data promises to facilitate a better understanding of the relat…

Cited by 0SourcecodeScholar
2025

Learning-Based Dynamic Robot-to-Human Handover

ICRA 2025

This paper presents a novel learning-based approach to dynamic robot-to-human handover, addressing the challenges of delivering objects to a moving receiver. We hypothesize that dynamic handover, where the robot adjusts to the receiver's movements, results in more efficient and comfortable interacti

Cited by 2SourcecodeScholar
2025

Wave-U-Mamba: An End-To-End Framework For High-Quality And Efficient Speech Super Resolution

ICASSP 2025accepted

Speech Super-Resolution (SSR) is a task of enhancing low-resolution speech signals by restoring missing high-frequency components. Conventional approaches typically reconstruct log-mel features, followed by a vocoder that generates high-resolution speech in the waveform domain. However, as mel featu…

Cited by 0SourceScholar
2024

AdaMER-CTC: Connectionist Temporal Classification with Adaptive Maximum Entropy Regularization for Automatic Speech Recognition

ICASSP 2024accepted

In Automatic Speech Recognition (ASR) systems, a recurring obstacle is the generation of narrowly focused output distributions. This phenomenon emerges as a side effect of Connectionist Temporal Classification (CTC), a robust sequence learning tool that utilizes dynamic programming for sequence mapp…

Cited by 0SourceScholar
2024

Class-Wise Buffer Management for Incremental Object Detection: An Effective Buffer Training Strategy

ICASSP 2024accepted

Class incremental learning aims to solve a problem that arises when continuously adding unseen class instances to an existing model This approach has been extensively studied in the context of image classification; however its applicability to object detection is not well established yet. Existing f…

Cited by 0SourceScholar
2024

Data Driven Grapheme-to-Phoneme Representations for a Lexicon-Free Text-to-Speech

ICASSP 2024accepted

Grapheme-to-Phoneme (G2P) is an essential first step in any modern, high-quality Text-to-Speech (TTS) system. Most of the current G2P systems rely on carefully hand-crafted lexicons developed by experts. This poses a two-fold problem. Firstly, the lexicons are generated using a fixed phoneme set, us…

Cited by 0SourceScholar
2024

Latent Filling: Latent Space Data Augmentation for Zero-Shot Speech Synthesis

ICASSP 2024accepted

Previous works in zero-shot text-to-speech (ZS-TTS) have attempted to enhance its systems by enlarging the training data through crowd-sourcing or augmenting existing speech data. However, the use of low-quality data has led to a decline in the overall system performance. To avoid such degradation,…

Cited by 0SourceScholar
2024

Mels-Tts : Multi-Emotion Multi-Lingual Multi-Speaker Text-To-Speech System Via Disentangled Style Tokens

ICASSP 2024accepted

This paper proposes a multi-emotion, multi-lingual, and multi-speaker text-to-speech (MELS-TTS) system, employing disentangled style tokens for effective emotion transfer. In speech encompassing various attributes, such as emotional state, speaker identity, and linguistic style, disentangling these…

Cited by 0SourceScholar
2024

Stochastic Amortization: A Unified Approach to Accelerate Feature and Data Attribution

NeurIPS 2024poster

Many tasks in explainable machine learning, such as data valuation and feature attribution, perform expensive computation for each data point and are intractable for large datasets. These methods require efficient approximations, and although amortizing the process by learning a network to directly…

Cited by 6SourcePDFScholar
2023

Contrastive Corpus Attribution for Explaining Representations

ICLR 2023poster

Despite the widespread use of unsupervised models, very few methods are designed to explain them. Most explanation methods explain a scalar model output. However, unsupervised models output representation vectors, the elements of which are not good candidates to explain because they lack semantic me…

2023

Counterfactual Two-Stage Debiasing For Video Corpus Moment Retrieval

ICASSP 2023accepted

Video Corpus Moment Retrieval aims to select a temporal video moment pertinent to a given language query from a large video corpus. Existing systems are prone to rely on a retrieval bias as a shortcut, which hinders the systems from accurately learning vision-language association. The retrieval bias…

Cited by 0SourceScholar
2023

Learning to Estimate Shapley Values with Vision Transformers

ICLR 2023top-25%

Transformers have become a default architecture in computer vision, but understanding what drives their predictions remains a challenging problem. Current explanation approaches rely on attention values or input gradients, but these provide a limited view of a model’s dependencies. Shapley values of…

2023

Self-Supervised Accent Learning for Under-Resourced Accents Using Native Language Data

ICASSP 2023accepted

In this paper, we propose a novel method to improve the accuracy of an English speech recognizer for a target accent using the corresponding native language data. Collecting labeled data for all accents of English to train an end-to-end neural speech recognizer for English is a difficult and expensi…

Cited by 0SourceScholar
2023

Transformer-Based Unified Recognition of Two Hands Manipulating Objects

CVPR 2023poster

Understanding the hand-object interactions from an egocentric video has received a great attention recently. So far, most approaches are based on the convolutional neural network (CNN) features combined with the temporal encoding via the long short-term memory (LSTM) or graph convolution network (GC…

2021

Neural Utterance Confidence Measure for RNN-Transducers and Two Pass Models

ICASSP 2021accepted

In this paper, we propose methods to compute confidence score on the predictions made by an end-to-end speech recognition model in a 2-pass framework. We use RNN-Transducer for a streaming model, and an attention-based decoder for the second pass model. We use neural technique to compute the confide…

Cited by 0SourceScholar
2021

Streaming End-to-End Speech Recognition with Jointly Trained Neural Feature Enhancement

ICASSP 2021accepted

In this paper, we present a streaming end-to-end speech recognition model based on Monotonic Chunkwise Attention (MoCha) jointly trained with enhancement layers. Even though the MoCha attention enables streaming speech recognition with recognition accuracy comparable to a full attention-based approa…

Cited by 0SourceScholar
2021

Task Aware Multi-Task Learning for Speech to Text Tasks

ICASSP 2021accepted

In general, the direct Speech-to-text translation (ST) is jointly trained with Automatic Speech Recognition (ASR), and Machine Translation (MT) tasks. However, the issues with the current joint learning strategies inhibit the knowledge transfer across these tasks. We propose a task modulation networ…

Cited by 0SourceScholar
2020

End-end Speech-to-Text Translation with Modality Agnostic Meta-Learning

ICASSP 2020accepted

Collecting large amounts of data to train end-to-end Speech Translation (ST) models is more difficult compared to the ASR and MT tasks. Previous studies have proposed the use of transfer learning approaches to overcome the above difficulty. These approaches benefit from weakly supervised training da…

Cited by 0SourceScholar
2020

Small Energy Masking for Improved Neural Network Training for End-To-End Speech Recognition

ICASSP 2020accepted

In this paper, we present a Small Energy Masking (SEM) algorithm, which masks inputs having values below a certain threshold. More specifically, a time-frequency bin is masked if the filterbank energy in this bin is less than a certain energy threshold. A uniform distribution is employed to randomly…

Cited by 0SourceScholar
2019

Robust Recognition of Reverberant and Noisy Speech Using Coherence-based Processing

ICASSP 2019accepted

This paper describes a combination of techniques for improving speech recognition accuracy using two microphones in reverberant and noisy environments. These techniques include both monaural and binaural processing. The first stage is monaural precedence-based processing that enhances the onsets of…

Cited by 0SourceScholar
2018

Sound Source Separation Using Phase Difference and Reliable Mask Selection Selection

ICASSP 2018accepted

In this paper, we present an algorithm called Reliable Mask Selection-Phase Difference Channel Weighting (RMS-PDCW) which selects the target source masked by a noise source using the Angle of Arrival (AoA) information calculated using the phase difference information. The RMS-PDCW algorithm selects…

Cited by 0SourceScholar
2018

Spectral Distortion Model for Training Phase-Sensitive Deep-Neural Networks for Far-Field Speech Recognition

ICASSP 2018accepted

In this paper, we present an algorithm which introduces phase-perturbation to the training database when training phase-sensitive deep neural-network models. Traditional features such as log-mel or cepstral features do not have have any phase-relevant information. However features such as raw-wavefo…

Cited by 3SourceScholar