← Search

Sumin Lee

12 accepted papers

2026

AC-Sampler: Accelerate and Correct Diffusion Sampling with Metropolis-Hastings Algorithm

ICLR 2026poster

Diffusion-based generative models have recently achieved state-of-the-art performance in high-fidelity image synthesis. These models learn a sequence of denoising transition kernels that gradually transform a simple prior distribution into a complex data distribution. However, requiring many transit…

Cited by 0SourcecodeScholar
2026

Generalizable Slum Detection from Satellite Imagery with Mixture-of-Experts

AAAI 2026technical

Satellite-based slum segmentation holds significant promise in generating global estimates of urban poverty. However, the morphological heterogeneity of informal settlements presents a major challenge, hindering the ability of models trained on specific regions to generalize effectively to unseen lo

Cited by 0SourcePDFScholar
2026

Geometric Embedding Alignment via Curvature Matching in Transfer Learning

ICML 2026poster

Geometrical interpretations of deep learning models offer insightful perspectives into their underlying mathematical structures. In this work, we introduce a novel approach that leverages differential geometry, particularly concepts from Riemannian geometry, to integrate multiple models into a unifi…

Cited by 0SourceScholar
2026

ModuLoop: Low-Level Code Generation Using Modular Synthesizer and Closed-Loop Debugger for Robotic Control

ICRA 2026poster

Large Language Models (LLMs) have demonstrated impressive performance across various domains, including code generation and problem solving. However, their application in robotic control—particularly in low-level tasks that require precise manipulation, real-time feedback, and environment-dependent …

2025

ModuLoop: Low-Level Code Generation Using Modular Synthesizer and Closed-Loop Debugger for Robotic Control

RA-L 2025

Large Language Models (LLMs) have demonstrated impressive performance across various domains, including code generation and problem solving. However, their application in robotic control—particularly in low-level tasks that require precise manipulation, real-time feedback, and environment-dependent

Cited by 0SourceScholar
2025

Score-informed Neural Operator for Enhancing Ordering-based Causal Discovery

NeurIPS 2025poster

Ordering-based approaches to causal discovery identify topological orders of causal graphs, providing scalable alternatives to combinatorial search methods. Under the Additive Noise Models (ANMs) assumption, recent causal ordering methods based on score matching require an accurate estimation of the…

Cited by 0SourceScholar
2025

Trajectory-Class-Aware Multi-Agent Reinforcement Learning

ICLR 2025poster

In the context of multi-agent reinforcement learning, *generalization* is a challenge to solve various tasks that may require different joint policies or coordination without relying on policies specialized for each task. We refer to this type of problem as a *multi-task*, and we train agents to be…

2024

Flow-Assisted Motion Learning Network for Weakly-Supervised Group Activity Recognition

ECCV 2024poster

"Weakly-Supervised Group Activity Recognition (WSGAR) aims to understand the activity performed together by a group of individuals with the video-level label and without actor-level labels. We propose Flow-Assisted Motion Learning Network () for WSGAR, which consists of the motion-aware actor encode…

Cited by 1SourcePDFScholar
2024

Geometrically Aligned Transfer Encoder for Inductive Transfer in Regression Tasks

ICLR 2024poster

Transfer learning is a crucial technique for handling a small amount of data that is potentially related to other abundant data. However, most of the existing methods are focused on classification tasks using images and language datasets. Therefore, in order to expand the transfer learning scheme to…

Cited by 2SourcePDFScholar
2023

Towards Good Practices for Missing Modality Robust Action Recognition

AAAI 2023technical

Standard multi-modal models assume the use of the same modalities in training and inference stages. However, in practice, the environment in which multi-modal models operate may not satisfy such assumption. As such, their performances degrade drastically if any modality is missing in the inference s…

2020

Hi-CMD: Hierarchical Cross-Modality Disentanglement for Visible-Infrared Person Re-Identification

CVPR 2020poster

Visible-infrared person re-identification (VI-ReID) is an important task in night-time surveillance applications, since visible cameras are difficult to capture valid appearance information under poor illumination conditions. Compared to traditional person re-identification that handles only the int…

Cited by 412PDFcodeScholar