← Search

Minseo Kim

10 accepted papers

2026

ConceptPrism: Concept Disentanglement in Personalized Diffusion Models via Residual Token Optimization

CVPR 2026

Personalized text-to-image (T2I) generation has emerged as a key application for creating user-specific concepts from a few reference images. The core challenge is concept disentanglement: separating the target concept from irrelevant residual information. Lacking such disentanglement, capturing hig

Cited by 0SourceScholar
2026

Stable-GFlowNet: Toward Diverse and Robust LLM Red-Teaming via Contrastive Trajectory Balance

ICML 2026spotlight

Large Language Model Red-Teaming, which proactively identifies vulnerabilities of large language models, is an essential process for ensuring safety. Finding effective and diverse attacks in red team activities is important, but achieving both is challenging. Generative Flow Networks (GFN) that perf…

Cited by 0SourceScholar
2025

Beyond Task-Oriented and Chitchat Dialogues: Proactive and Transition-Aware Conversational Agents

EMNLP 2025

Conversational agents have traditionally been developed for either task-oriented dialogue (TOD) or open-ended chitchat, with limited progress in unifying the two. Yet, real-world conversations naturally involve fluid transitions between these modes. To address this gap, we introduce TACT (TOD-And-Ch

2025

CANVAS: Commonsense-Aware Navigation System for Intuitive Human-Robot Interaction

ICRA 2025

Real-life robot navigation involves more than just reaching a destination; it requires optimizing movements while addressing scenario-specific goals. An intuitive way for humans to express these goals is through abstract cues like verbal commands or rough sketches. Such human guidance may lack detai

Cited by 4SourceScholar
2025

DisCoRD: Discrete Tokens to Continuous Motion via Rectified Flow Decoding

ICCV 2025poster

Human motion is inherently continuous and dynamic, posing significant challenges for generative models. While discrete generation methods are widely used, they suffer from limited expressiveness and frame-wise noise artifacts. In contrast, continuous approaches produce smoother, more natural motion…

Cited by 0SourcePDFScholar
2025

KRETA: A Benchmark for Korean Reading and Reasoning in Text-Rich VQA Attuned to Diverse Visual Contexts

EMNLP 2025

Understanding and reasoning over text within visual contexts poses a significant challenge for Vision-Language Models (VLMs), given the complexity and diversity of real-world scenarios. To address this challenge, text-rich Visual Question Answering (VQA) datasets and benchmarks have emerged for high

2025

Subtle Risks, Critical Failures: A Framework for Diagnosing Physical Safety of LLMs for Embodied Decision Making

EMNLP 2025

Large Language Models (LLMs) are increasingly used for decision making in embodied agents, yet existing safety evaluations often rely on coarse success rates and domain-specific setups, making it difficult to diagnose why and where these models fail. This obscures our understanding of embodied safet

Cited by 0SourcePDFScholar
2024

Compositional Video Understanding with Spatiotemporal Structure-based Transformers

CVPR 2024poster

In this paper we suggest a new novel method to understand complex semantic structures through long video inputs. Conventional methods for understanding videos have been focused on short-term clips and trained to get visual representations for the short clips using convolutional neural networks or tr…

2024

Selective Vision is the Challenge for Visual Reasoning: A Benchmark for Visual Argument Understanding

EMNLP 2024main

Visual arguments, often used in advertising or social causes, rely on images to persuade viewers to do or believe something. Understanding these arguments requires selective vision: only specific visual stimuli within an image are relevant to the argument, and relevance can only be understood within…