← Search

Kyungsu Kim

17 accepted papers

2026

GH-NAF: Grid-Adaptive Hash-Level-Attended Neural Attenuation Fields for Discrepancy-Aware CBCT

CVPR 2026

Neural radiance fields (NeRF)-based methods with multi-resolution hash encoding enable efficient sparse-view CBCT reconstruction, but real-world projections violate ideal assumptions due to scatter/noise and related inconsistencies. Uniformly fusing hash-grid levels entangles heterogeneous frequency

Cited by 0SourcecodeScholar
2026

On the Collapse of Generative Paths: A Criterion and Correction for Diffusion Steering

ICML 2026poster

Inference-time steering adapts pretrained diffusion and flow models to new tasks without retraining, often utilizing ratio-of-densities constructions that reweight time-indexed marginals with fixed exponents. We identify Marginal Path Collapse, a failure mode in which the intermediate density define…

Cited by 0SourceScholar
2026

TRACE: Your Diffusion Model is Secretly an Instance Edge Detector

ICLR 2026oral

High-quality instance and panoptic segmentation has traditionally relied on dense instance-level annotations such as masks, boxes, or points, which are costly, inconsistent, and difficult to scale. Unsupervised and weakly-supervised approaches reduce this burden but remain constrained by semantic ba…

Cited by 0SourcecodeScholar
2025

COIN: Confidence Score-Guided Distillation for Annotation-Free Cell Segmentation

ICCV 2025poster

Cell instance segmentation (CIS) is crucial for identifying individual cell morphologies in histopathological images, providing valuable insights for biological and medical research. While unsupervised CIS (UCIS) models aim to reduce the heavy reliance on labor-intensive image annotations, they fail…

2025

Early Timestep Zero-Shot Candidate Selection for Instruction-Guided Image Editing

ICCV 2025poster

Despite recent advances in diffusion models, achieving reliable image generation and editing results remains challenging due to the inherent diversity induced by stochastic noise in the sampling process. Particularly, instruction-guided image editing with diffusion models offers user-friendly editin…

2025

Token-Supervised Value Models for Enhancing Mathematical Problem-Solving Capabilities of Large Language Models

ICLR 2025poster

With the rapid advancement of test-time compute search strategies to improve the mathematical problem-solving capabilities of large language models (LLMs), the need for building robust verifiers has become increasingly important. However, all these inference strategies rely on existing verifiers ori…

Cited by 0SourcePDFScholar
2025

TokenSynth: A Token-based Neural Synthesizer for Instrument Cloning and Text-to-Instrument

ICASSP 2025accepted

Recent advancements in neural audio codecs have enabled the use of tokenized audio representations in various audio generation tasks, such as text-to-speech, text-to-audio, and text-to-music generation. Leveraging this approach, we propose TokenSynth, a novel neural synthesizer that utilizes a decod…

Cited by 0SourceScholar
2024

Data-Efficient Unsupervised Interpolation Without Any Intermediate Frame for 4D Medical Images

CVPR 2024poster

4D medical images which represent 3D images with temporal information are crucial in clinical practice for capturing dynamic changes and monitoring long-term disease progression. However acquiring 4D medical images poses challenges due to factors such as radiation exposure and imaging duration neces…

2023

GEX: A flexible method for approximating influence via Geometric Ensemble

NeurIPS 2023poster

Through a deeper understanding of predictions of neural networks, Influence Function (IF) has been applied to various tasks such as detecting and relabeling mislabeled samples, dataset pruning, and separation of data sources in practice. However, we found standard approximations of IF suffer from pe…

2023

MARS: Model-agnostic Biased Object Removal without Additional Supervision for Weakly-Supervised Semantic Segmentation

ICCV 2023poster

Weakly-supervised semantic segmentation aims to reduce labeling costs by training semantic segmentation models using weak supervision, such as image-level class labels. However, most approaches struggle to produce accurate localization maps and suffer from false predictions in class-related backgrou…

Cited by 27PDFcodeScholar
2023

Show Me the Instruments: Musical Instrument Retrieval From Mixture Audio

ICASSP 2023accepted

As digital music production has become mainstream, the selection of appropriate virtual instruments plays a crucial role in determining the quality of music. To search the musical instrument samples or virtual instruments that make one’s desired sound, music producers use their ears to listen and co…

Cited by 0SourceScholar
2020

BinaryDuo: Reducing Gradient Mismatch in Binary Activation Network by Coupling Binary Activations

ICLR 2020poster

Binary Neural Networks (BNNs) have been garnering interest thanks to their compute cost reduction and memory savings. However, BNNs suffer from performance degradation mainly due to the gradient mismatch caused by binarizing activations. Previous works tried to address the gradient mismatch problem…

Cited by 53SourcecodeScholar
2020

Modality Shifting Attention Network for Multi-Modal Video Question Answering

CVPR 2020poster

This paper considers a network referred to as Modality Shifting Attention Network (MSAN) for Multimodal Video Question Answering (MVQA) task. MSAN decomposes the task into two sub-tasks: (1) localization of temporal moment relevant to the question, and (2) accurate prediction of the answer based on…

Cited by 102PDFScholar
2020

Unifying Activation- and Timing-based Learning Rules for Spiking Neural Networks

NeurIPS 2020poster

For the gradient computation across the time domain in Spiking Neural Networks (SNNs) training, two different approaches have been independently studied. The first is to compute the gradients with respect to the change in spike activation (activation-based methods), and the second is to compute the…

2019

Learning Not to Learn: Training Deep Neural Networks With Biased Data

CVPR 2019poster

We propose a novel regularization algorithm to train deep neural networks, in which data at training time is severely biased. Since a neural network efficiently learns data distribution, a network is likely to learn the bias information to categorize input data. It leads to poor performance at test…

Cited by 526PDFScholar
2019

Progressive Attention Memory Network for Movie Story Question Answering

CVPR 2019poster

This paper proposes the progressive attention memory network (PAMN) for movie story question answering (QA). Movie story QA is challenging compared to VQA in two aspects: (1) pinpointing the temporal parts relevant to answer the question is difficult as the movies are typically longer than an hour,…

Cited by 96PDFScholar