← Search

Minhyuk Sung

44 accepted papers

2026

BézierFlow: Learning Bézier Stochastic Interpolant Schedulers for Few-Step Generation

ICLR 2026poster

We introduce BézierFlow, a lightweight training approach for few-step generation with pretrained diffusion and flow models. BézierFlow achieves a 2–3× performance improvement for sampling with $\leq$ 10 NFEs while requiring only 15 minutes of training. Recent lightweight training approaches have sho…

Cited by 0SourcecodeScholar
2026

PairFlow: Closed-Form Source-Target Coupling for Few-Step Generation in Discrete Flow Models

ICLR 2026poster

We introduce $\texttt{PairFlow}$, a lightweight preprocessing step for training Discrete Flow Models (DFMs) to achieve few-step sampling without requiring a pretrained teacher. DFMs have recently emerged as a new class of generative models for discrete data, offering strong performance. However, the…

Cited by 0SourcecodeScholar
2026

Projected Gradient Ascent for Efficient Reward-Guided Updates with One-Step Generative Models

ICML 2026poster

We propose a constrained latent optimization method for reward-guided generation that preserves white Gaussian noise characteristics with negligible overhead. Test-time latent optimization can unlock substantially better reward-guided generations from pretrained generative models, but it is prone to…

Cited by 0SourceScholar
2026

Token Warping Helps MLLMs Look from Nearby Viewpoints

CVPR 2026

Can warping tokens, rather than pixels, help multimodal large language models (MLLMs) understand how a scene appears from a nearby viewpoint? While MLLMs perform well on visual reasoning, they remain fragile to viewpoint changes, as pixel-wise warping is highly sensitive to small depth errors and of

Cited by 0SourcecodeScholar
2025

$\Psi$-Sampler: Initial Particle Sampling for SMC-Based Inference-Time Reward Alignment in Score Models

NeurIPS 2025spotlight

We introduce $\Psi$-Sampler, an SMC-based framework incorporating pCNL-based initial particle sampling for effective inference-time reward alignment with a score-based model. Inference-time reward alignment with score-based generative models has recently gained significant traction, following a broa…

Cited by 0SourceScholar
2025

Inference-Time Scaling for Flow Models via Stochastic Generation and Rollover Budget Forcing

NeurIPS 2025poster

We propose an inference-time scaling approach for pretrained flow models. Recently, inference-time scaling has gained significant attention in LLMs and diffusion models, improving sample quality or better aligning outputs with user preferences by leveraging additional computation. For diffusion mode…

Cited by 0SourceScholar
2025

Moment- and Power-Spectrum-Based Gaussianity Regularization for Text-to-Image Models

NeurIPS 2025poster

We propose a novel regularization loss that enforces standard Gaussianity, encouraging samples to align with a standard Gaussian distribution. This facilitates a range of downstream tasks involving optimization in the latent space of text-to-image models. We treat elements of a high-dimensional samp…

Cited by 0SourceScholar
2025

ORIGEN: Zero-Shot 3D Orientation Grounding in Text-to-Image Generation

NeurIPS 2025poster

We introduce ORIGEN, the first zero-shot method for 3D orientation grounding in text-to-image generation across multiple objects and diverse categories. While previous work on spatial grounding in image generation has mainly focused on 2D positioning, it lacks control over 3D orientation. To address…

Cited by 0SourceScholar
2025

Perspective-Aware Reasoning in Vision-Language Models via Mental Imagery Simulation

ICCV 2025poster

We present a framework for perspective-aware reasoning in vision-language models (VLMs) through mental imagery simulation. Perspective-taking, the ability to perceive an environment or situation from an alternative viewpoint, is a key benchmark for human-level visual understanding, essential for env…

Cited by 0SourcePDFScholar
2025

REWIND: Real-Time Egocentric Whole-Body Motion Diffusion with Exemplar-Based Identity Conditioning

CVPR 2025poster

We present REWIND (Real-Time Egocentric Whole-Body Motion Diffusion), a one-step diffusion model for real-time, high-fidelity human motion estimation from egocentric image inputs. While an existing method for egocentric whole-body (i.e., body and hands) motion estimation is non-real-time and acausal…

Cited by 0SourcePDFScholar
2025

StochSync: Stochastic Diffusion Synchronization for Image Generation in Arbitrary Spaces

ICLR 2025poster

We propose a zero-shot method for generating images in arbitrary spaces (e.g., a sphere for 360◦ panoramas and a mesh surface for texture) using a pretrained image diffusion model. The zero-shot generation of various visual content using a pretrained image diffusion model has been explored mainly in…

Cited by 0SourcePDFScholar
2025

VideoHandles: Editing 3D Object Compositions in Videos Using Video Generative Priors

CVPR 2025poster

Generative methods for image and video editing use generative models as priors to perform edits despite incomplete information, such as changing the composition of 3D objects shown in a single image. Recent methods have shown promising composition editing results in the image setting, but in the vid…

Cited by 1SourcePDFScholar
2024

As-Plausible-As-Possible: Plausibility-Aware Mesh Deformation Using 2D Diffusion Priors

CVPR 2024poster

We present As-Plausible-as-Possible (APAP) mesh deformation technique that leverages 2D diffusion priors to preserve the plausibility of a mesh under user-controlled deformation. Our framework uses per-face Jacobians to represent mesh deformations where mesh vertex coordinates are computed via a dif…

2024

GrounDiT: Grounding Diffusion Transformers via Noisy Patch Transplantation

NeurIPS 2024poster

We introduce GrounDiT, a novel training-free spatial grounding technique for text-to-image generation using Diffusion Transformers (DiT). Spatial grounding with bounding boxes has gained attention for its simplicity and versatility, allowing for enhanced user control in image generation. However, pr…

2024

InterHandGen: Two-Hand Interaction Generation via Cascaded Reverse Diffusion

CVPR 2024poster

We present InterHandGen a novel framework that learns the generative prior of two-hand interaction. Sampling from our model yields plausible and diverse two-hand shapes in close interaction with or without an object. Our prior can be incorporated into any optimization or learning methods to reduce a…

Cited by 13SourcePDFScholar
2024

MV2Cyl: Reconstructing 3D Extrusion Cylinders from Multi-View Images

NeurIPS 2024poster

We present MV2Cyl, a novel method for reconstructing 3D from 2D multi-view images, not merely as a field or raw geometry but as a sketch-extrude CAD. Extracting extrusion cylinders from raw 3D geometry has been extensively researched in computer vision, while the processing of 3D data through neural…

Cited by 1SourcePDFScholar
2024

Neural Pose Representation Learning for Generating and Transferring Non-Rigid Object Poses

NeurIPS 2024poster

We propose a novel method for learning representations of poses for 3D deformable objects, which specializes in 1) disentangling pose information from the object's identity, 2) facilitating the learning of pose variations, and 3) transferring pose information to other object identities. Based on the…

Cited by 0SourcePDFScholar
2024

SyncTweedies: A General Generative Framework Based on Synchronized Diffusions

NeurIPS 2024poster

We introduce a general diffusion synchronization framework for generating diverse visual content, including ambiguous images, panorama images, 3D mesh textures, and 3D Gaussian splats textures, using a pretrained image diffusion model. We first present an analysis of various scenarios for synchroniz…

2023

FourierHandFlow: Neural 4D Hand Representation Using Fourier Query Flow

NeurIPS 2023poster

Recent 4D shape representations model continuous temporal evolution of implicit shapes by (1) learning query flows without leveraging shape and articulation priors or (2) decoding shape occupancies separately for each time value. Thus, they do not effectively capture implicit correspondences between…

Cited by 5SourcePDFScholar
2023

Im2Hands: Learning Attentive Implicit Representation of Interacting Two-Hand Shapes

CVPR 2023poster

We present Implicit Two Hands (Im2Hands), the first neural implicit representation of two interacting hands. Unlike existing methods on two-hand reconstruction that rely on a parametric hand model and/or low-resolution meshes, Im2Hands can produce fine-grained geometry of two hands with high hand-to…

2023

SALAD: Part-Level Latent Diffusion for 3D Shape Generation and Manipulation

ICCV 2023poster

We present a cascaded diffusion model based on a part-level implicit 3D representation. Our model achieves state-of-the-art generation quality and also enables part-level shape editing and manipulation without any additional training in conditional setup. Diffusion models have demonstrated impressiv…

Cited by 49PDFScholar
2023

ShapeTalk: A Language Dataset and Framework for 3D Shape Edits and Deformations

CVPR 2023poster

Editing 3D geometry is a challenging task requiring specialized skills. In this work, we aim to facilitate the task of editing the geometry of 3D models through the use of natural language. For example, we may want to modify a 3D chair model to "make its legs thinner" or to "open a hole in its back"…

2023

SyncDiffusion: Coherent Montage via Synchronized Joint Diffusions

NeurIPS 2023poster

The remarkable capabilities of pretrained image diffusion models have been utilized not only for generating fixed-size images but also for creating panoramas. However, naive stitching of multiple images often results in visible seams. Recent techniques have attempted to address this issue by perform…

2022

Implicit LiDAR Network: LiDAR Super-Resolution via Interpolation Weight Prediction

ICRA 2022poster

Super-resolution of LiDAR range images is crucial to improving many downstream tasks such as object detection, recognition, and tracking. While deep learning has made a remarkable advances in super-resolution techniques, typical convolutional architectures limit upscaling factors to specific output…

Cited by 33SourceScholar
2022

LADIS: Language Disentanglement for 3D Shape Editing

EMNLP 2022finding

Natural language interaction is a promising direction for democratizing 3D shape design. However, existing methods for text-driven 3D shape editing face challenges in producing decoupled, local edits to 3D shapes. We address this problem by learning disentangled latent representations that ground la…

2022

PartGlot: Learning Shape Part Segmentation From Language Reference Games

CVPR 2022oral

We introduce PartGlot, a neural framework and associated architectures for learning semantic part segmentation of 3D shape geometry, based solely on part referential language. We exploit the fact that linguistic descriptions of a shape can provide priors on the shape's parts -- as natural language h…

Cited by 33PDFcodeScholar
2022

Point2Cyl: Reverse Engineering 3D Objects From Point Clouds to Extrusion Cylinders

CVPR 2022poster

We propose Point2Cyl, a supervised network transforming a raw 3D point cloud to a set of extrusion cylinders. Reverse engineering from a raw geometry to a CAD model is an essential task to enable manipulation of the 3D data in shape editing software and thus expand their usages in many downstream ap…

Cited by 64PDFScholar
2022

Pop-Out Motion: 3D-Aware Image Deformation via Learning the Shape Laplacian

CVPR 2022poster

We propose a framework that can deform an object in a 2D image as it exists in 3D space. Most existing methods for 3D-aware image manipulation are limited to (1) only changing the global scene information or depth, or (2) manipulating an object of specific categories. In this paper, we present a 3D-…

Cited by 3PDFScholar
2022

The Shape Part Slot Machine: Contact-Based Reasoning for Generating 3D Shapes from Parts

ECCV 2022poster

"We present the Shape Part Slot Machine, a new method for assembling novel 3D shapes from existing parts by performing contact-based reasoning. Our method represents each shape as a graph of ""slots,"" where each slot is a region of contact between two shape parts. Based on this representation, we d…

Cited by 12SourcePDFScholar
2021

CPFN: Cascaded Primitive Fitting Networks for High-Resolution Point Clouds

ICCV 2021poster

Representing human-made objects as a collection of base primitives has a long history in computer vision and reverse engineering. In the case of high-resolution point cloud scans, the challenge is to be able to detect both large primitives as well as those explaining the detailed parts. While the cl…

Cited by 38PDFcodeScholar
2021

CTRL-C: Camera Calibration TRansformer With Line-Classification

ICCV 2021poster

Single image camera calibration is the task of estimating the camera parameters from a single input image, such as the vanishing points, focal length, and horizon line. In this work, we propose Camera calibration TRansformer with Line-Classification (CTRL-C), an end-to-end neural network-based appro…

Cited by 49PDFcodeScholar
2021

DeepMetaHandles: Learning Deformation Meta-Handles of 3D Meshes With Biharmonic Coordinates

CVPR 2021poster

We propose DeepMetaHandles, a 3D conditional generative model based on mesh deformation. Given a collection of 3D meshes of a category and their deformation handles (control points), our method learns a set of meta-handles for each shape, which are represented as combinations of the given handles. T…

Cited by 41PDFcodeScholar
2021

Joint Learning of 3D Shape Retrieval and Deformation

CVPR 2021poster

We propose a novel technique for producing high-quality 3D models that match a given target object image or scan. Our method is based on retrieving an existing shape from a database of 3D models and then deforming its parts to match the target shape. Unlike previous approaches that independently foc…

Cited by 48PDFScholar
2021

MultiBodySync: Multi-Body Segmentation and Motion Estimation via 3D Scan Synchronization

CVPR 2021poster

We present MultiBodySync, a novel, end-to-end trainable multi-body motion segmentation and rigid registration framework for multiple input 3D point clouds. The two non-trivial challenges posed by this multi-scan multibody setting that we investigate are: (i) guaranteeing correspondence and segmentat…

Cited by 57PDFcodeScholar
2020

Deformation-Aware 3D Model Embedding and Retrieval

ECCV 2020poster

We introduce a new problem of mph{retrieving} 3D models that are mph{deformable} to a given query shape and present a novel deep mph{deformation-aware} embedding to solve this retrieval task. 3D model retrieval is a fundamental operation for recovering a clean and complete 3D model from a noisy and…

2020

Learning 3D Part Assembly from a Single Image

ECCV 2020poster

Autonomous assembly is a crucial capability for robots in many applications. For this task, several problems such as obstacle avoidance, motion planning, and actuator control have been extensively studied in robotics. However, when it comes to task specification, the space of possibilities remains u…

2020

Neural Geometric Parser for Single Image Camera Calibration

ECCV 2020poster

We propose a neural geometric parser learning single image camera calibration for man-made scenes. Unlike previous neural approaches that rely only on semantic cues obtained from neural networks, our approach considers both semantic and geometric cues, resulting in significant accuracy improvement.…

2020

Pix2Surf: Learning Parametric 3D Surface Models of Objects from Images

ECCV 2020poster

We investigate the problem of learning to generate 3D parametric surface representations for novel object instances, as seen from one or more views. Previous work on learning shape reconstruction from multiple views uses discrete representations such as point clouds or voxels, while continuous surfa…

Cited by 42SourcePDFScholar
2019

GSPN: Generative Shape Proposal Network for 3D Instance Segmentation in Point Cloud

CVPR 2019poster

We introduce a novel 3D object proposal approach named Generative Shape Proposal Network (GSPN) for instance segmentation in point cloud data. Instead of treating object proposal as a direct bounding box regression problem, we take an analysis-by-synthesis strategy and generate proposals by reconstr…

Cited by 381PDFcodeScholar
2019

Supervised Fitting of Geometric Primitives to 3D Point Clouds

CVPR 2019oral

Fitting geometric primitives to 3D point cloud data bridges a gap between low-level digitized 3D data and high-level structural information on the underlying 3D shapes. As such, it enables many downstream applications in 3D data processing. For a long time, RANSAC-based methods have been the gold st…

Cited by 246PDFScholar
2018

Deep Functional Dictionaries: Learning Consistent Semantic Structures on 3D Models from Functions

NeurIPS 2018poster

Various 3D semantic attributes such as segmentation masks, geometric features, keypoints, and materials can be encoded as per-point probe functions on 3D geometries. Given a collection of related 3D shapes, we consider how to jointly analyze such probe functions over different shapes, and how to dis…