← Search

Minkyu Kim

19 accepted papers

2026

Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games

ICLR 2026poster

Large Language Model (LLM) agents are reshaping the game industry, by enabling more intelligent and human-preferable characters. Yet, current game benchmarks fall short of practical needs: they lack evaluations of diverse LLM capabilities across various game genres, studies of agentic modules crucia…

Cited by 0SourceScholar
2026

See and Fix the Flaws: Enabling VLMs and Diffusion Models to Comprehend Visual Artifacts via Agentic Data Synthesis

CVPR 2026

Despite recent advances in diffusion models, AI generated images still often contain visual artifacts that compromise realism. Although more thorough pre-training and bigger models might reduce artifacts, there is no assurance that they can be completely eliminated, which makes artifact mitigation a

Cited by 0SourcecodeScholar
2026

VLM-SubtleBench: How Far Are VLMs from Human-Level Subtle Comparative Reasoning?

ICLR 2026poster

The ability to distinguish subtle differences between visually similar images is essential for diverse domains such as industrial anomaly detection, medical imaging, and aerial surveillance. While comparative reasoning benchmarks for vision-language models (VLMs) have recently emerged, they primaril…

Cited by 0SourcecodeScholar
2025

Energy-based generator matching: A neural sampler for general state space

NeurIPS 2025poster

We propose Energy-based generator matching (EGM), a modality-agnostic approach to train generative models from energy functions in the absence of data. Extending the recently proposed generator matching, EGM enables training of arbitrary continuous-time Markov processes, e.g., diffusion, flow, and j…

Cited by 0SourceScholar
2025

On scalable and efficient training of diffusion samplers

NeurIPS 2025poster

We address the challenge of training diffusion models to sample from unnormalized energy distributions in the absence of data, the so-called diffusion samplers. Although these approaches have shown promise, they struggle to scale in more demanding scenarios where energy evaluations are expensive and…

Cited by 0SourceScholar
2025

Rare-to-Frequent: Unlocking Compositional Generation Power of Diffusion Models on Rare Concepts with LLM Guidance

ICLR 2025spotlight

State-of-the-art text-to-image (T2I) diffusion models often struggle to generate rare compositions of concepts, e.g., objects with unusual attributes. In this paper, we show that the compositional generation power of diffusion models on such rare concepts can be significantly enhanced by the Large L…

2025

Test-time Alignment of Diffusion Models without Reward Over-optimization

ICLR 2025spotlight

Diffusion models excel in generative tasks, but aligning them with specific objectives while maintaining their versatility remains challenging. Existing fine-tuning methods often suffer from reward over-optimization, while approximate guidance approaches fail to optimize target rewards effectively.…

2024

Discovering and Mitigating Visual Biases through Keyword Explanation

CVPR 2024highlight

Addressing biases in computer vision models is crucial for real-world AI deployments. However mitigating visual biases is challenging due to their unexplainable nature often identified indirectly through visualization or sample statistics which necessitates additional human supervision for interpret…

2024

Image Clustering Conditioned on Text Criteria

ICLR 2024poster

Classical clustering methods do not provide users with direct control of the clustering results, and the clustering results may not be consistent with the relevant criterion that a user has in mind. In this work, we present a new methodology for performing image clustering based on user-specified cr…

2024

Neural Image Compression with Text-guided Encoding for both Pixel-level and Perceptual Fidelity

ICML 2024poster

Recent advances in text-guided image compression have shown great potential to enhance the perceptual quality of reconstructed images. These methods, however, tend to have significantly degraded pixel-wise fidelity, limiting their practicality. To fill this gap, we develop a new text-guided image co…

2023

Balanced Column-Wise Block Pruning for Maximizing GPU Parallelism

AAAI 2023technical

Pruning has been an effective solution to reduce the number of computations and the memory requirement in deep learning. The pruning unit plays an important role in exploiting the GPU resources efficiently. The filter is proposed as a simple pruning unit of structured pruning. However, since the fi…

2023

S-CLIP: Semi-supervised Vision-Language Learning using Few Specialist Captions

NeurIPS 2023poster

Vision-language models, such as contrastive language-image pre-training (CLIP), have demonstrated impressive results in natural image domains. However, these models often struggle when applied to specialized domains like remote sensing, and adapting to such domains is challenging due to the limited…

2021

SmoothMix: Training Confidence-calibrated Smoothed Classifiers for Certified Robustness

NeurIPS 2021poster

Randomized smoothing is currently a state-of-the-art method to construct a certifiably robust classifier from neural networks against $\ell_2$-adversarial perturbations. Under the paradigm, the robustness of a classifier is aligned with the prediction confidence, i.e., the higher confidence from a s…

2019

Toward Achieving Formal Guarantees for Human-Aware Controllers in Human-Robot Interactions

IROS 2019poster

With the primary objective of human-robot interaction being to support humans' goals, there exists a need to formally synthesize robot controllers that can provide the desired service. Synthesis techniques have the benefit of providing formal guarantees for specification satisfaction. There is poten…

Cited by 8SourceScholar
2016

Tele-operation system with reliable grasping force estimation to compensate for the time-varying sEMG feature

ICRA 2016

This paper presents a real-time framework for tele-manipulation by using sEMG signals to estimate both human motion and force intention. Our previous study showed that the ability to detect discrete force levels was not applicable to complex tasks such as grasping, holding, and manipulating various

Cited by 11SourceScholar
2015

A robust control method of multi-DOF power-assistant robots for unknown external perturbation using sEMG signals

IROS 2015poster

This paper presents a control method of multi-DOF power assistant robots for anatomical multi-axis joints such as the wrist and the ankle. It is difficult to calculate the accurate direction of human motion intention during manipulating an object due to discrepancy between the calculated force from…

Cited by 8SourceScholar