← Search

Young Min Kim

34 accepted papers

2026

3D-aware Disentangled Representation for Compositional Reinforcement Learning

ICLR 2026poster

Vision-based reinforcement learning can benefit from object-centric scene representation, which factorizes the visual observation into individual objects and their attributes, such as color, shape, size, and position. While such object-centric representations can extract components that generalize w…

Cited by 0SourceScholar
2026

Point2Act: Efficient 3D Distillation of Multimodal LLMs for Zero-Shot Context-Aware Grasping

ICRA 2026poster

We propose Point2Act, which directly retrieves the 3D action point relevant to a contextually described task, leveraging Multimodal Large Language Models (MLLMs). Foundation models have opened the possibility for generalist robots that can perform a zero-shot task following natural language descript…

2025

Event-Driven Storytelling with Multiple Lifelike Humans in a 3D Scene

ICCV 2025poster

In this work, we propose a framework that creates a lively virtual dynamic scene with contextual motions of multiple humans. Generating multi-human contextual motion requires holistic reasoning over dynamic relationships among human-human and human-scene interactions. We adapt the power of a large l…

Cited by 0SourcePDFScholar
2025

Humans as a Calibration Pattern: Dynamic 3D Scene Reconstruction from Unsynchronized and Uncalibrated Videos

ICCV 2025poster

Recent works on dynamic 3D neural field reconstruction assume the input from synchronized multi-view videos whose poses are known. The input constraints are often not satisfied in real-world setups, making the approach impractical. We show that unsynchronized videos from unknown poses can generate d…

Cited by 0SourcePDFScholar
2025

Isometric Regularization for Manifolds of Functional Data

ICLR 2025poster

While conventional data are represented as discrete vectors, Implicit Neural Representations (INRs) utilize neural networks to represent data points as continuous functions. By incorporating a shared network that maps latent vectors to individual functions, one can model the distribution of function…

Cited by 0SourcePDFScholar
2025

Learning 3D Scene Analogies with Neural Contextual Scene Maps

ICCV 2025poster

Understanding scene contexts is crucial for machines to perform tasks and adapt prior knowledge in unseen or noisy 3D environments. As data-driven learning is intractable to comprehensively encapsulate diverse ranges of layouts and open spaces, we propose teaching machines to identify relational com…

2025

Less is More: Improving Motion Diffusion Models with Sparse Keyframes

ICCV 2025poster

Recent advances in motion diffusion models have led to remarkable progress in diverse motion generation tasks, including text-to-motion synthesis.However, existing approaches represent motions as dense frame sequences, requiring the model to process redundant or less informative frames.The processin…

Cited by 0SourcePDFScholar
2025

SceneMI: Motion In-betweening for Modeling Human-Scene Interaction

ICCV 2025poster

Modeling human-scene interactions (HSI) is essential for understanding and simulating everyday human behaviors. Recent approaches utilizing generative modeling have made progress in this domain; however, they are limited in controllability and flexibility for real-world applications. To address thes…

Cited by 0SourcePDFScholar
2024

Outdoor Scene Extrapolation with Hierarchical Generative Cellular Automata

CVPR 2024highlight

We aim to generate fine-grained 3D geometry from large-scale sparse LiDAR scans abundantly captured by autonomous vehicles (AV). Contrary to prior work on AV scene completion we aim to extrapolate fine geometry from unlabeled and beyond spatial limits of LiDAR scans taking a step towards generating…

Cited by 0SourcePDFScholar
2023

Calibrating Panoramic Depth Estimation for Practical Localization and Mapping

ICCV 2023poster

The absolute depth values of surrounding environments provide crucial cues for various assistive technologies, such as localization, navigation, and 3D structure estimation. We propose that accurate depth estimated from panoramic images can serve as a powerful and light-weight input for a wide range…

Cited by 2PDFcodeScholar
2023

FindAdaptNet: Find and Insert Adapters by Learned Layer Importance

ICASSP 2023accepted

Adapters are lightweight bottleneck modules introduced to assist pre-trained self-supervised learning (SSL) models to be customized to new tasks. However, searching the appropriate layers to insert adapters on large models has become difficult due to the large number of possible layers and thus a va…

Cited by 0SourceScholar
2023

Text2Scene: Text-Driven Indoor Scene Stylization With Part-Aware Details

CVPR 2023highlight

We propose Text2Scene, a method to automatically create realistic textures for virtual scenes composed of multiple objects. Guided by a reference image and text descriptions, our pipeline adds detailed texture on labeled 3D geometries in the room such that the generated colors respect the hierarchic…

Cited by 15SourcePDFScholar
2022

CPO: Change Robust Panorama to Point Cloud Localization

ECCV 2022poster

"We present CPO, a fast and robust algorithm that localizes a 2D panorama with respect to a 3D point cloud of a scene possibly containing changes. To robustly handle scene changes, our approach deviates from conventional feature point matching, and focuses on the spatial context provided from panora…

2022

MasKGrasp: Mask-based Grasping for Scenes with Multiple General Real-world Objects

IROS 2022poster

In this paper, we introduce a mask-based grasping method that discerns multiple objects within the scene regard-less of transparency or specularity and finds the optimal grasp position avoiding clutter. Conventional vision-based robotic grasping approaches often fail to extend to the scenes containi…

Cited by 3SourceScholar
2022

MoDA: Map Style Transfer for Self-Supervised Domain Adaptation of Embodied Agents

ECCV 2022poster

"We propose a domain adaptation method, MoDA, which adapts a pretrained embodied agent to a new, noisy environment without ground-truth supervision. Map-based memory provides important contextual information for visual navigation, and exhibits unique spatial structure mainly composed of flat walls a…

Cited by 10SourcePDFScholar
2022

Neural Marionette: Unsupervised Learning of Motion Skeleton and Latent Dynamics from Volumetric Video

AAAI 2022technical

We present Neural Marionette, an unsupervised approach that discovers the skeletal structure from a dynamic sequence and learns to generate diverse motions that are consistent with the observed motion dynamics. Given a video stream of point cloud observation of an articulated body under arbitrary mo…

Cited by 6SourcePDFScholar
2021

GATSBI: Generative Agent-Centric Spatio-Temporal Object Interaction

CVPR 2021poster

We present GATSBI, a generative model that can transform a sequence of raw observations into a structured latent representation that fully captures the spatio-temporal context of the agent's actions. In vision-based decision-making scenarios, an agent faces complex high-dimensional observations wher…

Cited by 7PDFcodeScholar
2021

Learning to Generate 3D Shapes with Generative Cellular Automata

ICLR 2021poster

In this work, we present a probabilistic 3D generative model, named Generative Cellular Automata, which is able to produce diverse and high quality shapes. We formulate the shape generation process as sampling from the transition kernel of a Markov chain, where the sampling chain eventually evolves…

Cited by 30SourcePDFScholar
2021

N-ImageNet: Towards Robust, Fine-Grained Object Recognition With Event Cameras

ICCV 2021poster

We introduce N-ImageNet, a large-scale dataset targeted for robust, fine-grained object recognition with event cameras. The dataset is collected using programmable hardware in which an event camera consistently moves around a monitor displaying images from ImageNet. N-ImageNet serves as a challengin…

Cited by 105PDFcodeScholar
2021

Structure-From-Sherds: Incremental 3D Reassembly of Axially Symmetric Pots From Unordered and Mixed Fragment Collections

ICCV 2021poster

Re-assembling multiple pots accurately from numerous 3D scanned fragments remains a challenging task to this date. Previous methods extract all potential matching pairs of pot sherds and considers them simultaneously to search for an optimal global pot configuration. In this work, we empirically sho…

Cited by 12PDFScholar
2019

RL-GAN-Net: A Reinforcement Learning Agent Controlled GAN Network for Real-Time Point Cloud Shape Completion

CVPR 2019poster

We present RL-GAN-Net, where a reinforcement learning (RL) agent provides fast and robust control of a generative adversarial network (GAN). Our framework is applied to point cloud shape completion that converts noisy, partial point cloud data into a high-fidelity completed shape by controlling the…

Cited by 243PDFcodeScholar