← Search

Soohyun Kim

14 accepted papers

2026

Mitigating Hallucination in Vision-Language Model with Depth and Spatial-aware Key-Value Refinement

ICLR 2026poster

Large vision–language models (VLMs) deliver state-of-the-art results on a wide range of multimodal tasks, yet they remain prone to visual hallucinations, producing content that is not grounded in the input image. Despite progress with visual supervision, reinforcement learning, and post-hoc attenti…

Cited by 0SourceScholar
2025

Subtractive Training for Music Stem Insertion Using Latent Diffusion Models

ICASSP 2025accepted

We present Subtractive Training<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup>, a simple and novel method for synthesizing individual musical instrument stems given other instruments as context. This method pairs a dataset of complete music mixe…

Cited by 0SourceScholar
2024

Diffusion-driven GAN Inversion for Multi-Modal Face Image Generation

CVPR 2024poster

We present a new multi-modal face image generation method that converts a text prompt and a visual input such as a semantic mask or scribble map into a photo-realistic face image. To do this we combine the strengths of Generative Adversarial networks (GANs) and diffusion models (DMs) by employing th…

2023

LANIT: Language-Driven Image-to-Image Translation for Unlabeled Data

CVPR 2023poster

Existing techniques for image-to-image translation commonly have suffered from two critical problems: heavy reliance on per-sample domain annotation and/or inability to handle multiple attributes per image. Recent truly-unsupervised methods adopt clustering approaches to easily provide per-sample on…

2023

Robust Camera Pose Refinement for Multi-Resolution Hash Encoding

ICML 2023poster

Multi-resolution hash encoding has recently been proposed to reduce the computational cost of neural renderings, such as NeRF. This method requires accurate camera poses for the neural renderings of given scenes. However, contrary to previous methods jointly optimizing camera poses and 3D scenes, th…

Cited by 28SourcePDFScholar
2022

Deep Translation Prior: Test-Time Training for Photorealistic Style Transfer

AAAI 2022technical

Recent techniques to solve photorealistic style transfer within deep convolutional neural networks (CNNs) generally require intensive training from large-scale datasets, thus having limited applicability and poor generalization ability to unseen images or styles. To overcome this, we propose a novel…

2022

InstaFormer: Instance-Aware Image-to-Image Translation With Transformer

CVPR 2022poster

We present a novel Transformer-based network architecture for instance-aware image-to-image translation, dubbed InstaFormer, to effectively integrate global- and instance-level information. By considering extracted content features from an image as tokens, our networks discover global consensus of c…

Cited by 65PDFcodeScholar
2021

Correlate-and-Excite: Real-Time Stereo Matching via Guided Cost Volume Excitation

IROS 2021poster

Volumetric deep learning approach towards stereo matching aggregates a cost volume computed from input left and right images using 3D convolutions. Recent works showed that utilization of extracted image features and a spatially varying cost volume aggregation complements 3D convolutions. However, e…

Cited by 85SourcecodeScholar
2017

Designing Anthropomorphic Robot Hand With Active Dual-Mode Twisted String Actuation Mechanism and Tiny Tension Sensors

RA-L 2017

In this letter, using the active dual-mode twisted string actuation (TSA) mechanism and tiny tension sensors on the tendon strings, an anthropomorphic robot hand is newly designed in a compact manner. Thanks to the active dual-mode TSA mechanism, which is a miniaturized transmission, the proposed ro

Cited by 96SourceScholar
2017

Modified nonlinear pressure estimator of pneumatic actuator for force controller design

ICRA 2017poster

This paper presents a modified nonlinear pneumatic model for pneumatic force servo systems. The modified model is proposed in order to estimate pressures accurately for both chambers of a pneumatic cylinder by adopting flow coefficient maps, which is different from a conventional model whose flow co…

Cited by 1SourceScholar
2016

Two-channel electrotactile stimulation for sensory feedback of fingers of prosthesis

IROS 2016poster

Electrotactile stimulation has been used to provide sensory information of forearm prosthesis to users. Although conventional sensory feedback method, where one electrode expresses sensory information of only one finger, could provide force information of three fingers by using three electrodes, it…

Cited by 25SourceScholar
2015

Dual-mode twisting actuation mechanism with an active clutch for active mode-change and simple relaxation process

IROS 2015poster

In this paper, a dual-mode twisting actuation mechanism with an active clutch is newly presented for a high performance tendon-driven robot (e.g., robot hand). This mechanism is a kind of mechanical automatic power transmission mechanism which provides fast motion and large contraction force by two…

Cited by 10SourceScholar
2015

The SoftGait: A simple and powerful weight-support device for walking and squatting

IROS 2015poster

When designing a lower-limb assisting robot with body-weight support (BWS), it is important to achieve high force fidelity to support body weight at standing phase. Low impedance operation at swing phase is also required not to disturb leg motions for users. The SoftGait is designed to achieve both…

Cited by 18SourceScholar
2015

Using common spatial pattern algorithm for unsupervised real-time estimation of fingertip forces from sEMG signals

IROS 2015poster

In this paper, a method to extract the fingertip forces of the index and middle fingers from surface electromyography (sEMG) signals is studied by adopting the known common spatial pattern (CSP) approach. For unsupervised estimation of fingertip forces in real-time, CSP filtering is shown to be a no…

Cited by 11SourceScholar