← Search

Minsu Kim

75 accepted papers

2026

Active Attacks: Red-teaming LLMs via Adaptive Environments

ICML 2026poster

We address the challenge of automatically generating diverse attack prompts for large language models (LLMs) that elicit harmful behaviors (e.g., insults, sexual content) and are used for safety fine-tuning. While several prior approaches train LLMs with reinforcement learning (RL) to generate such …

Cited by 0SourceScholar
2026

Diffusion Alignment as Variataional Expectation-Maximization

ICLR 2026poster

Diffusion alignment aims to optimize diffusion models for the downstream objective. While existing methods based on reinforcement learning or direct backpropagation achieve considerable success in maximizing rewards, they often suffer from reward over-optimization and mode collapse. We introduce Dif…

Cited by 0SourcecodeScholar
2026

IMSE: Intrinsic Mixture of Spectral Experts Fine-tuning for Test-Time Adaptation

ICLR 2026poster

Test-time adaptation (TTA) has been widely explored to prevent performance degradation when test data differ from the training distribution. However, fully leveraging the rich representations of large pretrained models with minimal parameter updates remains underexplored. In this paper, we propose a…

Cited by 0SourcecodeScholar
2026

Inlier-Centric Post-Training Quantization for Object Detection Models

ICLR 2026poster

Object detection is pivotal in robotics, but its immense computational demands make the models slow and power-hungry, underscoring the need for quantization. However, when the quantization is applied in practice, cluttered backgrounds and irregular object morphologies cause redundant activations (or…

Cited by 0SourceScholar
2026

Latent Veracity Inference for Identifying Errors in Stepwise Reasoning

ICLR 2026poster

Chain-of-Thought (CoT) reasoning has advanced the capabilities and transparency of language models (LMs); however, reasoning chains can contain inaccurate statements that reduce performance and trustworthiness. To address this, we propose to augment each reasoning step in a CoT with a latent veracit…

Cited by 0SourceScholar
2026

Synthesizable Molecular Generation via Soft-constrained GFlowNets with Rich Chemical Priors

ICML 2026poster

The application of generative models for experimental drug discovery campaigns is severely limited by the difficulty of designing molecules de novo that can be synthesized in practice. Previous works have leveraged Generative Flow Networks (GFlowNets) to impose hard synthesizability constraints thro…

Cited by 0SourceScholar
2026

UniSpector: Towards Universal Open-set Defect Recognition via Spectral-Contrastive Visual Prompting

CVPR 2026

Although industrial inspection systems should be capable of recognizing unprecedented defects, most existing approaches operate under a closed-set assumption, which prevents them from detecting novel anomalies. While visual prompting offers a scalable alternative for industrial inspection, existing

Cited by 0SourcecodeScholar
2025

A Deep Reinforcement Learning based End-to-End Control Framework for Lower Limb Exoskeletons with Smooth Movement Transitions

IROS 2025

This paper presents an active control strategy for lower limb exoskeletons by proposing an end-to-end framework employing deep reinforcement learning (DRL) to enable smooth transitions between different movement patterns. The majority of existing methods in exoskeleton literature employ finite state

Cited by 0SourceScholar
2025

Adaptive Inference-Time Scaling via Cyclic Diffusion Search

NeurIPS 2025poster

Diffusion models have demonstrated strong generative capabilities across domains ranging from image synthesis to complex reasoning tasks. However, most inference-time scaling methods rely on fixed denoising schedules, limiting their ability to allocate computation based on instance difficulty or tas…

Cited by 0SourceScholar
2025

Adaptive teachers for amortized samplers

ICLR 2025poster

Amortized inference is the task of training a parametric model, such as a neural network, to approximate a distribution with a given unnormalized density where exact sampling is intractable. When sampling is modeled as a sequential decision-making process, reinforcement learning (RL) methods, such a…

2025

Ant Colony Sampling with GFlowNets for Combinatorial Optimization

AISTATS 2025poster

We present the Generative Flow Ant Colony Sampler (GFACS), a novel meta-heuristic method that hierarchically combines amortized inference and parallel stochastic search. Our method first leverages Generative Flow Networks (GFlowNets) to amortize a multi-modal prior distribution over combinatorial so…

Cited by 0SourceScholar
2025

Contextual Speech Extraction: Leveraging Textual History as an Implicit Cue for Target Speech Extraction

ICASSP 2025accepted

In this paper, we investigate a novel approach for Target Speech Extraction (TSE), which relies solely on textual context to extract the target speech. We refer to this task as Contextual Speech Extraction (CSE). Unlike traditional TSE methods that rely on pre-recorded enrollment utterances, video o…

Cited by 0SourceScholar
2025

Energy-based generator matching: A neural sampler for general state space

NeurIPS 2025poster

We propose Energy-based generator matching (EGM), a modality-agnostic approach to train generative models from energy functions in the absence of data. Extending the recently proposed generator matching, EGM enables training of arbitrary continuous-time Markov processes, e.g., diffusion, flow, and j…

Cited by 0SourceScholar
2025

ExploreGS: Explorable 3D Scene Reconstruction with Virtual Camera Samplings and Diffusion Priors

ICCV 2025poster

Recent advances in novel view synthesis (NVS) have enabled real-time rendering with 3D Gaussian Splatting (3DGS). However, existing methods struggle with artifacts and missing regions when rendering unseen viewpoints, limiting seamless scene exploration. To address this, we propose a 3DGS-based pipe…

Cited by 0SourcePDFScholar
2025

From Evidence to Belief: A Bayesian Epistemology Approach to Language Models

NAACL 2025long

This paper investigates the knowledge of language models from the perspective of Bayesian epistemology. We explore how language models adjust their confidence and responses when presented with evidence with varying levels of informativeness and reliability. To study these properties, we create a dat…

Cited by 0SourcePDFScholar
2025

Generative Flows on Synthetic Pathway for Drug Design

ICLR 2025poster

Generative models in drug discovery have recently gained attention as efficient alternatives to brute-force virtual screening. However, most existing models do not account for synthesizability, limiting their practical use in real-world scenarios. In this paper, we propose RxnFlow, which sequentiall…

2025

Improved Off-policy Reinforcement Learning in Biological Sequence Design

ICML 2025poster

Designing biological sequences with desired properties is challenging due to vast search spaces and limited evaluation budgets. Although reinforcement learning methods use proxy models for rapid reward evaluation, insufficient training data can cause proxy misspecification on out-of-distribution inp…

2025

Large Language Models are Strong Audio-Visual Speech Recognition Learners

ICASSP 2025accepted

Multimodal large language models (MLLMs) have recently become a focal point of research due to their formidable multimodal understanding capabilities. For example, in the audio and speech domains, an LLM can be equipped with (automatic) speech recognition (ASR) abilities by just concatenating the au…

Cited by 0SourceScholar
2025

Learning Diverse Attacks on Large Language Models for Robust Red-Teaming and Safety Tuning

ICLR 2025poster

Red-teaming, or identifying prompts that elicit harmful responses, is a critical step in ensuring the safe and responsible deployment of large language models (LLMs). Developing effective protection against many modes of attack prompts requires discovering diverse attacks. Automated red-teaming typi…

2025

MOFFlow: Flow Matching for Structure Prediction of Metal-Organic Frameworks

ICLR 2025poster

Metal-organic frameworks (MOFs) are a class of crystalline materials with promising applications in many areas such as carbon capture and drug delivery. In this work, we introduce MOFFlow, the first deep generative model tailored for MOF structure prediction. Existing approaches, including ab initio…

2025

MoME: Mixture of Matryoshka Experts for Audio-Visual Speech Recognition

NeurIPS 2025poster

Large language models (LLMs) have recently shown strong potential in audio-visual speech recognition (AVSR), but their high computational demands and sensitivity to token granularity limit their practicality in resource-constrained settings. Token compression methods can reduce inference cost, but t…

Cited by 0SourceScholar
2025

On scalable and efficient training of diffusion samplers

NeurIPS 2025poster

We address the challenge of training diffusion models to sample from unnormalized energy distributions in the absence of data, the so-called diffusion samplers. Although these approaches have shown promise, they struggle to scale in more demanding scenarios where energy evaluations are expensive and…

Cited by 0SourceScholar
2025

Outsourced Diffusion Sampling: Efficient Posterior Inference in Latent Spaces of Generative Models

ICML 2025poster

Any well-behaved generative model over a variable $\mathbf{x}$ can be expressed as a deterministic transformation of an exogenous (‘*outsourced'*) Gaussian noise variable $\mathbf{z}$: $\mathbf{x}=f_\theta(\mathbf{z})$. In such a model (*eg*, a VAE, GAN, or continuous-time flow-based model), sampli…

Cited by 0SourcePDFScholar
2025

T-CIL: Temperature Scaling using Adversarial Perturbation for Calibration in Class-Incremental Learning

CVPR 2025poster

We study model confidence calibration in class-incremental learning, where models learn from sequential tasks with different class sets. While existing works primarily focus on accuracy, maintaining calibrated confidence has been largely overlooked. Unfortunately, most post-hoc calibration technique…

Cited by 0SourcePDFScholar
2025

Trajectory Balance with Asynchrony: Decoupling Exploration and Learning for Fast, Scalable LLM Post-Training

NeurIPS 2025poster

Reinforcement learning (RL) is a critical component of large language model (LLM) post-training. However, on-policy algorithms used for post-training are not naturally robust to a diversified content of experience replay buffers, which asynchronous off-policy actors can efficiently populate in paral…

Cited by 0SourcecodeScholar
2025

Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations

ICCV 2025poster

We explore a novel zero-shot Audio-Visual Speech Recognition (AVSR) framework, dubbed Zero-AVSR, which enables speech recognition in target languages without requiring any audio-visual speech data in those languages. Specifically, we introduce the Audio-Visual Speech Romanizer (AV-Romanizer), which…

2024

AV2AV: Direct Audio-Visual Speech to Audio-Visual Speech Translation with Unified Audio-Visual Speech Representation

CVPR 2024highlight

This paper proposes a novel direct Audio-Visual Speech to Audio-Visual Speech Translation (AV2AV) framework where the input and output of the system are multimodal (i.e. audio and visual speech). With the proposed AV2AV two key advantages can be brought: 1) We can perform real-like conversations wit…

2024

Addressing Diverging Training Costs Using BEVRestore for High-Resolution Bird's Eye View Map Construction

RA-L 2024

Recent advancements in Bird's Eye View (BEV) fusion for map construction have demonstrated remarkable mapping of urban environments. However, their deep and bulky architecture incurs substantial amounts of backpropagation memory and computing latency. Consequently, the problem poses an unavoidable b

Cited by 1SourceScholar
2024

Amortizing intractable inference in diffusion models for vision, language, and control

NeurIPS 2024poster

Diffusion models have emerged as effective distribution estimators in vision, language, and reinforcement learning, but their use as priors in downstream tasks poses an intractable posterior inference problem. This paper studies *amortized* sampling of the posterior over data, $\mathbf{x}\sim p^{\rm…

2024

Analysis of the Memorization and Generalization Capabilities of AI Agents: are Continual Learners Robust?

ICASSP 2024accepted

In continual learning (CL), an AI agent (e.g., autonomous vehicles or robotics) learns from non-stationary data streams under dynamic environments. For the practical deployment of such applications, it is important to guarantee robustness to unseen environments while maintaining past experiences. In…

Cited by 0SourceScholar
2024

BroadBEV: Collaborative LiDAR-camera Fusion for Broad-sighted Bird’s Eye View Map Construction

ICRA 2024poster

A recent sensor fusion in a Bird’s Eye View (BEV) space has shown its utility in various tasks such as 3D detection, map segmentation, etc. However, the approach struggles with inaccurate camera BEV estimation, and a perception of distant areas due to the sparsity of LiDAR points. In this paper, we…

Cited by 5SourceScholar
2024

Equity-Transformer: Solving NP-Hard Min-Max Routing Problems as Sequential Generation with Equity Context

AAAI 2024technical

Min-max routing problems aim to minimize the maximum tour length among multiple agents as they collaboratively visit all cities, i.e., the completion time. These problems include impactful real-world applications but are known as NP-hard. Existing methods are facing challenges, particularly in large…

2024

Exploring Phonetic Context-Aware Lip-Sync for Talking Face Generation

ICASSP 2024accepted

Talking face generation is the challenging task of synthesizing a natural and realistic face that requires accurate synchronization with a given audio. Due to co-articulation, where an isolated phone is influenced by the preceding or following phones, the articulation of a phone varies upon the phon…

Cited by 0SourceScholar
2024

Genetic-guided GFlowNets for Sample Efficient Molecular Optimization

NeurIPS 2024poster

The challenge of discovering new molecules with desired properties is crucial in domains like drug discovery and material design. Recent advances in deep learning-based generative methods have shown promise but face the issue of sample efficiency due to the computational expense of evaluating the re…

2024

Hierarchically Structured Neural Bones for Reconstructing Animatable Objects from Casual Videos

ECCV 2024poster

"We propose a new framework for creating and easily manipulating 3D models of arbitrary objects using casually captured videos. Our core ingredient is a novel hierarchy deformation model, which captures motions of objects with a tree-structured bones. Our hierarchy system decomposes motions based on…

2024

Improved off-policy training of diffusion samplers

NeurIPS 2024poster

We study the problem of training diffusion models to sample from a distribution with a given unnormalized density or energy function. We benchmark several diffusion-structured inference methods, including simulation-based variational approaches and off-policy methods (continuous generative flow netw…

2024

Learning to Scale Logits for Temperature-Conditional GFlowNets

ICML 2024poster

GFlowNets are probabilistic models that sequentially generate compositional structures through a stochastic policy. Among GFlowNets, temperature-conditional GFlowNets can introduce temperature-based controllability for exploration and exploitation. We propose *Logit-scaling GFlowNets* (Logit-GFN), a…

2024

Let’s Go Real Talk: Spoken Dialogue Model for Face-to-Face Conversation

ACL 2024long

In this paper, we introduce a novel Face-to-Face spoken dialogue model. It processes audio-visual speech from user input and generates audio-visual speech as the response, marking the initial step towards creating an avatar chatbot system without relying on intermediate text. To this end, we newly i…

2024

Local Search GFlowNets

ICLR 2024spotlight

Generative Flow Networks (GFlowNets) are amortized sampling methods that learn a distribution over discrete objects proportional to their rewards. GFlowNets exhibit a remarkable ability to generate diverse samples, yet occasionally struggle to consistently produce samples with high rewards due to ov…

2024

Pessimistic Backward Policy for GFlowNets

NeurIPS 2024poster

This paper studies Generative Flow Networks (GFlowNets), which learn to sample objects proportionally to a given reward function through the trajectory of state transitions. In this work, we observe that GFlowNets tend to under-exploit the high-reward objects due to training on insufficient number o…

2024

Quilt: Robust Data Segment Selection against Concept Drifts

AAAI 2024technical

Continuous machine learning pipelines are common in industrial settings where models are periodically trained on data streams. Unfortunately, concept drifts may occur in data streams where the joint distribution of the data X and label y, P(X, y), changes over time and possibly degrade model accurac…

Cited by 1SourcePDFScholar
2024

SpaFL: Communication-Efficient Federated Learning With Sparse Models And Low Computational Overhead

NeurIPS 2024poster

The large communication and computation overhead of federated learning (FL) is one of the main challenges facing its practical deployment over resource-constrained clients and systems. In this work, SpaFL: a communication-efficient FL framework is proposed to optimize sparse model structures with l…

2024

Symmetric Replay Training: Enhancing Sample Efficiency in Deep Reinforcement Learning for Combinatorial Optimization

ICML 2024poster

Deep reinforcement learning (DRL) has significantly advanced the field of combinatorial optimization (CO). However, its practicality is hindered by the necessity for a large number of reward evaluations, especially in scenarios involving computationally intensive function assessments. To enhance the…

2024

Text-Driven Talking Face Synthesis by Reprogramming Audio-Driven Models

ICASSP 2024accepted

In this paper, we present a method for reprogramming pre-trained audio-driven talking face synthesis models to operate in a text-driven manner. Consequently, we can easily generate face videos that articulate the provided textual sentences, eliminating the necessity of recording speech for each infe…

Cited by 0SourceScholar
2024

Towards Practical and Efficient Image-to-Speech Captioning with Vision-Language Pre-Training and Multi-Modal Tokens

ICASSP 2024accepted

In this paper, we propose methods to build a powerful and efficient Image-to-Speech captioning (Im2Sp) model. To this end, we start with importing the rich knowledge related to image comprehension and language modeling from a large-scale pre-trained vision-language model into Im2Sp. We set the outpu…

Cited by 0SourceScholar
2024

Visual Speech Recognition for Languages with Limited Labeled Data Using Automatic Labels from Whisper

ICASSP 2024accepted

This paper proposes a powerful Visual Speech Recognition (VSR) method for multiple languages, especially for low-resource languages that have a limited number of labeled data. Different from previous methods that tried to improve the VSR performance for the target language by using knowledge learned…

Cited by 0SourceScholar
2024

Where Visual Speech Meets Language: VSP-LLM Framework for Efficient and Context-Aware Visual Speech Processing

EMNLP 2024finding

In visual speech processing, context modeling capability is one of the most important requirements due to the ambiguous nature of lip movements. For example, homophenes, words that share identical lip movements but produce different sounds, can be distinguished by considering the context. In this pa…

2023

Bootstrapped Training of Score-Conditioned Generator for Offline Design of Biological Sequences

NeurIPS 2023poster

We study the problem of optimizing biological sequences, e.g., proteins, DNA, and RNA, to maximize a black-box score function that is only evaluated in an offline dataset. We propose a novel solution, bootstrapped training of score-conditioned generator (BootGen) algorithm. Our algorithm repeats a t…

2023

Deep Visual Forced Alignment: Learning to Align Transcription with Talking Face Video

AAAI 2023technical

Forced alignment refers to a technology that time-aligns a given transcription with a corresponding speech. However, as the forced alignment technologies have developed using speech audio, they might fail in alignment when the input speech audio is noise-corrupted or is not accessible. We focus on t…

Cited by 3SourcePDFScholar
2023

DevFormer: A Symmetric Transformer for Context-Aware Device Placement

ICML 2023poster

In this paper, we present DevFormer, a novel transformer-based architecture for addressing the complex and computationally demanding problem of hardware design optimization. Despite the demonstrated efficacy of transformers in domains including natural language processing and computer vision, their…

Cited by 21SourcePDFScholar
2023

Lip Reading for Low-resource Languages by Learning and Combining General Speech Knowledge and Language-specific Knowledge

ICCV 2023poster

This paper proposes a novel lip reading framework, especially for low-resource languages, which has not been well addressed in the previous literature. Since low-resource languages do not have enough video-text paired data to train the model to have sufficient power to model lip movements and langua…

Cited by 18PDFScholar
2023

Meta-SAGE: Scale Meta-Learning Scheduled Adaptation with Guided Exploration for Mitigating Scale Shift on Combinatorial Optimization

ICML 2023poster

This paper proposes Meta-SAGE, a novel approach for improving the scalability of deep reinforcement learning models for combinatorial optimization (CO) tasks. Our method adapts pre-trained models to larger-scale problems in test time by suggesting two components: a scale meta-learner (SML) and sched…

2023

PartMix: Regularization Strategy To Learn Part Discovery for Visible-Infrared Person Re-Identification

CVPR 2023poster

Modern data augmentation using a mixture-based technique can regularize the models from overfitting to the training data in various computer vision applications, but a proper data augmentation technique tailored for the part-based Visible-Infrared person Re-IDentification (VI-ReID) models remains un…

Cited by 93SourcePDFScholar
2023

Watch or Listen: Robust Audio-Visual Speech Recognition With Visual Corruption Modeling and Reliability Scoring

CVPR 2023poster

This paper deals with Audio-Visual Speech Recognition (AVSR) under multimodal input corruption situation where audio inputs and visual inputs are both corrupted, which is not well addressed in previous research directions. Previous studies have focused on how to complement the corrupted audio inputs…

2022

CERT: Continual Pre-training on Sketches for Library-oriented Code Generation

IJCAI 2022poster

Code generation is a longstanding challenge, aiming to generate a code snippet based on a natural language description. Usually, expensive text-code paired data is essential for training a code generation model. Recently, thanks to the success of pre-training techniques, large language models are tr…

2022

Distinguishing Homophenes Using Multi-Head Visual-Audio Memory for Lip Reading

AAAI 2022technical

Recognizing speech from silent lip movement, which is called lip reading, is a challenging task due to 1) the inherent information insufficiency of lip movement to fully represent the speech, and 2) the existence of homophenes that have similar lip movement with different pronunciations. In this pap…

Cited by 66SourcePDFScholar
2022

Sym-NCO: Leveraging Symmetricity for Neural Combinatorial Optimization

NeurIPS 2022accept

Deep reinforcement learning (DRL)-based combinatorial optimization (CO) methods (i.e., DRL-NCO) have shown significant merit over the conventional CO solvers as DRL-NCO is capable of learning CO solvers less relying on problem-specific expert domain knowledge (heuristic method) and supervised labele…

2022

SyncTalkFace: Talking Face Generation with Precise Lip-Syncing via Audio-Lip Memory

AAAI 2022technical

The challenge of talking face generation from speech lies in aligning two different modal information, audio and video, such that the mouth region corresponds to input audio. Previous methods either exploit audio-visual representation learning or leverage intermediate structural information such as…

Cited by 90SourcePDFScholar
2022

VisageSynTalk: Unseen Speaker Video-to-Speech Synthesis via Speech-Visage Feature Selection

ECCV 2022poster

"The goal of this work is to reconstruct speech from a silent talking face video. Recent studies have shown impressive performance on synthesizing speech from silent talking face videos. However, they have not explicitly considered on varying identity characteristics of different speakers, which pla…

Cited by 7SourcePDFScholar
2021

Cross-Domain Grouping and Alignment for Domain Adaptive Semantic Segmentation

AAAI 2021technical

Existing techniques to adapt semantic segmentation networks across source and target domains within deep convolutional neural networks (CNNs) deal with all the samples from the two domains in a global or category-aware manner. They do not consider an inter-class variation within the target domain it…

2021

Learning Canonical 3D Object Representation for Fine-Grained Recognition

ICCV 2021poster

We propose a novel framework for fine-grained object recognition that learns to recover object variation in 3D space from a single image, trained on an image collection without using any ground-truth 3D annotation. We accomplish this by representing an object as a composition of 3D shape and its app…

Cited by 16PDFScholar
2021

Learning Collaborative Policies to Solve NP-hard Routing Problems

NeurIPS 2021poster

Recently, deep reinforcement learning (DRL) frameworks have shown potential for solving NP-hard routing problems such as the traveling salesman problem (TSP) without problem-specific expert knowledge. Although DRL can be used to solve complex problems, DRL frameworks still struggle to compete with s…

2021

Multi-Modality Associative Bridging Through Memory: Speech Sound Recollected From Face Video

ICCV 2021poster

In this paper, we introduce a novel audio-visual multi-modal bridging framework that can utilize both audio and visual information, even with uni-modal inputs. We exploit a memory network that stores source (i.e., visual) and target (i.e., audio) modal representations, where source modal representat…

Cited by 52PDFScholar
2020

Cylindrical Convolutional Networks for Joint Object Detection and Viewpoint Estimation

CVPR 2020poster

Existing techniques to encode spatial invariance within deep convolutional neural networks only model 2D transformation fields. This does not account for the fact that objects in a 2D space are a projection of 3D ones, and thus they have limited ability to severe object viewpoint changes. To overcom…

Cited by 19PDFScholar
2020

Fast Adaptation of Deep Reinforcement Learning-Based Navigation Skills to Human Preference

ICRA 2020poster

Deep reinforcement learning (RL) is being actively studied for robot navigation due to its promise of superior performance and robustness. However, most existing deep RL navigation agents are trained using fixed parameters, such as maximum velocities and weightings of reward components. Since the op…

Cited by 25SourceScholar
2019

Deep Reinforcement Learning of Navigation in a Complex and Crowded Environment with a Limited Field of View

ICRA 2019poster

Mobile robots are required to navigate freely in a complex and crowded environment in order to provide services to humans. For this navigation ability, deep reinforcement learning (DRL)-based methods are gaining increasing attentions. However, existing DRL methods require a wide field of view (FOV),…

Cited by 85SourceScholar