← Search

Inwoo Hwang

17 accepted papers

2026

Towards Spatially Consistent Image Generation: On Incorporating Intrinsic Scene Properties into Diffusion Models

AAAI 2026technical

Image generation models trained on large datasets can synthesize high-quality images but often produce spatially inconsistent and distorted images due to limited information about the underlying structures and spatial layouts. In this work, we leverage intrinsic scene properties (e.g., depth, segmen

Cited by 0SourcePDFScholar
2025

Event-Driven Storytelling with Multiple Lifelike Humans in a 3D Scene

ICCV 2025poster

In this work, we propose a framework that creates a lively virtual dynamic scene with contextual motions of multiple humans. Generating multi-human contextual motion requires holistic reasoning over dynamic relationships among human-human and human-scene interactions. We adapt the power of a large l…

Cited by 0SourcePDFScholar
2025

Less is More: Improving Motion Diffusion Models with Sparse Keyframes

ICCV 2025poster

Recent advances in motion diffusion models have led to remarkable progress in diverse motion generation tasks, including text-to-motion synthesis.However, existing approaches represent motions as dense frame sequences, requiring the model to process redundant or less informative frames.The processin…

Cited by 0SourcePDFScholar
2025

PEER Pressure: Model-to-Model Regularization for Single Source Domain Generalization

CVPR 2025poster

Data augmentation is a popular tool for single source domain generalization, which expands the source domain by generating simulated ones, improving generalization on unseen target domains. In this work, we show that the performance of such augmentation-based methods in the target domains universall…

2025

SceneMI: Motion In-betweening for Modeling Human-Scene Interaction

ICCV 2025poster

Modeling human-scene interactions (HSI) is essential for understanding and simulating everyday human behaviors. Recent approaches utilizing generative modeling have made progress in this domain; however, they are limited in controllability and flexibility for real-world applications. To address thes…

Cited by 0SourcePDFScholar
2024

Efficient Monte Carlo Tree Search via On-the-Fly State-Conditioned Action Abstraction

UAI 2024poster

Monte Carlo Tree Search (MCTS) has showcased its efficacy across a broad spectrum of decision-making problems. However, its performance often degrades under vast combinatorial action space, especially where an action is composed of multiple sub-actions. In this work, we propose an action abstraction…

2024

Fine-Grained Causal Dynamics Learning with Quantization for Improving Robustness in Reinforcement Learning

ICML 2024poster

Causal dynamics learning has recently emerged as a promising approach to enhancing robustness in reinforcement learning (RL). Typically, the goal is to build a dynamics model that makes predictions based on the causal relationships among the entities. Despite the fact that causal connections often m…

2023

Learning Geometry-Aware Representations by Sketching

CVPR 2023poster

Understanding geometric concepts, such as distance and shape, is essential for understanding the real world and also for many vision tasks. To incorporate such information into a visual representation of a scene, we propose learning to represent the scene by sketching, inspired by human behavior. Ou…

Cited by 7SourcePDFScholar
2023

Text2Scene: Text-Driven Indoor Scene Stylization With Part-Aware Details

CVPR 2023highlight

We propose Text2Scene, a method to automatically create realistic textures for virtual scenes composed of multiple objects. Guided by a reference image and text descriptions, our pipeline adds detailed texture on labeled 3D geometries in the room such that the generated colors respect the hierarchic…

Cited by 15SourcePDFScholar
2022

MasKGrasp: Mask-based Grasping for Scenes with Multiple General Real-world Objects

IROS 2022poster

In this paper, we introduce a mask-based grasping method that discerns multiple objects within the scene regard-less of transparency or specularity and finds the optimal grasp position avoiding clutter. Conventional vision-based robotic grasping approaches often fail to extend to the scenes containi…

Cited by 3SourceScholar
2022

SelecMix: Debiased Learning by Contradicting-pair Sampling

NeurIPS 2022accept

Neural networks trained with ERM (empirical risk minimization) sometimes learn unintended decision rules, in particular when their training data is biased, i.e., when training labels are strongly correlated with undesirable features. To prevent a network from learning such features, recent methods a…