← Search

Jaesik Park

61 accepted papers

2026

Direct Reward Fine-Tuning on Poses for Single Image to 3D Human in the Wild

ICLR 2026poster

Single-view 3D human reconstruction has achieved remarkable progress through the adoption of multi-view diffusion models, yet the recovered 3D humans often exhibit unnatural poses. This phenomenon becomes pronounced when reconstructing 3D humans with dynamic or challenging poses, which we attribute…

Cited by 0SourceScholar
2026

MotionStream: Real-Time Video Generation with Interactive Motion Controls

ICLR 2026oral

Current motion-conditioned video generation methods suffer from prohibitive latency (minutes per video) and non-causal processing that prevents real-time interaction. We present MotionStream, enabling sub-second latency with up to 29 FPS streaming generation on a single GPU. Our approach begins by a…

Cited by 0SourcecodeScholar
2026

TRACE: Your Diffusion Model is Secretly an Instance Edge Detector

ICLR 2026oral

High-quality instance and panoptic segmentation has traditionally relied on dense instance-level annotations such as masks, boxes, or points, which are costly, inconsistent, and difficult to scale. Unsupervised and weakly-supervised approaches reduce this burden but remain constrained by semantic ba…

Cited by 0SourcecodeScholar
2025

BUFFER-X: Towards Zero-Shot Point Cloud Registration in Diverse Scenes

ICCV 2025poster

Recent advances in deep learning-based point cloud registration have improved generalization, yet most methods still require retraining or manual parameter tuning for each new environment. In this paper, we identify three key factors limiting generalization: (a) reliance on environment-specific voxe…

2025

Exploring Multimodal Diffusion Transformers for Enhanced Prompt-based Image Editing

ICCV 2025poster

Transformer-based diffusion models have recently superseded traditional U-Net architectures, with multimodal diffusion transformers (MM-DiT) emerging as the dominant approach in state-of-the-art models like Stable Diffusion 3 and Flux.1. Previous approaches have relied on unidirectional cross-attent…

Cited by 0SourcePDFScholar
2025

KISS-Matcher: Fast and Robust Point Cloud Registration Revisited

ICRA 2025

While global point cloud registration systems have advanced significantly in all aspects, many studies have focused on specific components, such as feature extraction, graph-theoretic pruning, or pose solvers. In this paper, we take a holistic view on the registration problem and develop an open-sou

Cited by 19SourcecodeScholar
2024

3D Geometric Shape Assembly via Efficient Point Cloud Matching

ICML 2024poster

Learning to assemble geometric shapes into a larger target structure is a pivotal task in various practical applications. In this work, we tackle this problem by establishing local correspondences between point clouds of part shapes in both coarse- and fine-levels. To this end, we introduce Proxy Ma…

2024

Distilling Diffusion Models into Conditional GANs

ECCV 2024poster

"We propose a method to distill a complex multistep diffusion model into a single-step conditional GAN student model, dramatically accelerating inference, while preserving image quality. Our approach interprets diffusion distillation as a paired image-to-image translation task, using noise-to-image…

Cited by 39SourcePDFScholar
2024

Extending CLIP’s Image-Text Alignment to Referring Image Segmentation

NAACL 2024long

Referring Image Segmentation (RIS) is a cross-modal task that aims to segment an instance described by a natural language expression. Recent methods leverage large-scale pretrained unimodal models as backbones along with fusion techniques for joint reasoning across modalities. However, the inherent…

Cited by 9SourcePDFScholar
2024

Learning SO(3)-Invariant Semantic Correspondence via Local Shape Transform

CVPR 2024poster

Establishing accurate 3D correspondences between shapes stands as a pivotal challenge with profound implications for computer vision and robotics. However existing self-supervised methods for this problem assume perfect input shape alignment restricting their real-world applicability. In this work w…

Cited by 2SourcePDFScholar
2024

Pick-or-Mix: Dynamic Channel Sampling for ConvNets

CVPR 2024poster

Channel pruning approaches for convolutional neural networks (ConvNets) deactivate the channels statically or dynamically and require special implementation. In addition channel squeezing in representative ConvNets is carried out via 1 x 1 convolutions which dominates a large portion of computations…

2023

Binary Radiance Fields

NeurIPS 2023poster

In this paper, we propose \textit{binary radiance fields} (BiRF), a storage-efficient radiance field representation employing binary feature encoding in a format of either $+1$ or $-1$. This binarization strategy lets us represent the feature grid with highly compact feature encoding and a dramatic…

Cited by 39SourcePDFScholar
2023

Scaling Up GANs for Text-to-Image Synthesis

CVPR 2023highlight

The recent success of text-to-image synthesis has taken the world by storm and captured the general public's imagination. From a technical standpoint, it also marked a drastic change in the favored architecture to design generative image models. GANs used to be the de facto choice, with techniques l…

Cited by 613SourcePDFScholar
2023

Scene-level Point Cloud Colorization with Semantics-and-geometry-aware Networks

ICRA 2023poster

In robotic applications, we often obtain tons of 3D point cloud data without color information, and it is difficult to visualize point clouds in a meaningful and colorful way. Can we colorize 3D point clouds for better visualization? Existing deep learning-based colorization methods usually only tak…

Cited by 3SourceScholar
2023

Spacetime Surface Regularization for Neural Dynamic Scene Reconstruction

ICCV 2023poster

We propose an algorithm, 4DRegSDF, for the spacetime surface regularization to improve the fidelity of neural rendering and reconstruction in dynamic scenes. The key idea is to impose local rigidity on the deformable Signed Distance Function (SDF) for temporal coherency. Our approach works by (1) sa…

Cited by 10PDFcodeScholar
2023

Stable and Consistent Prediction of 3D Characteristic Orientation via Invariant Residual Learning

ICML 2023poster

Learning to predict reliable characteristic orientations of 3D point clouds is an important yet challenging problem, as different point clouds of the same class may have largely varying appearances. In this work, we introduce a novel method to decouple the shape geometry and semantics of the input p…

Cited by 3SourcePDFScholar
2022

A Rotated Hyperbolic Wrapped Normal Distribution for Hierarchical Representation Learning

NeurIPS 2022accept

We present a rotated hyperbolic wrapped normal distribution (RoWN), a simple yet effective alteration of a hyperbolic wrapped normal distribution (HWN). The HWN expands the domain of probabilistic modeling from Euclidean to hyperbolic space, where a tree can be embedded with arbitrary low distortion…

2022

CostDCNet: Cost Volume Based Depth Completion for a Single RGB-D Image

ECCV 2022poster

"Successful depth completion from a single RGB-D image requires both extracting plentiful 2D and 3D features and merging these heterogeneous features appropriately. We propose a novel depth completion framework, CostDCNet, based on the cost volume-based depth estimation approach that has been succes…

2022

Learning Debiased Classifier with Biased Committee

NeurIPS 2022accept

Neural networks are prone to be biased towards spurious correlations between classes and latent attributes exhibited in a major portion of training data, which ruins their generalization capability. We propose a new method for training debiased classifiers with no spurious attribute label. The key i…

2022

PeRFception: Perception using Radiance Fields

NeurIPS 2022accept

The recent progress in implicit 3D representation, i.e., Neural Radiance Fields (NeRFs), has made accurate and photorealistic 3D reconstruction possible in a differentiable manner. This new representation can effectively convey the information of hundreds of high-resolution images in one compact for…

2022

PointMixer: MLP-Mixer for Point Cloud Understanding

ECCV 2022poster

"MLP-Mixer has newly appeared as a new challenger against the realm of CNNs and Transformer. Despite its simplicity compared to Transformer, the concept of channel-mixing MLPs and token-mixing MLPs achieves noticeable performance in image recognition tasks. Unlike images, point clouds are inherently…

2021

Brick-by-Brick: Combinatorial Construction with Deep Reinforcement Learning

NeurIPS 2021poster

Discovering a solution in a combinatorial space is prevalent in many real-world problems but it is also challenging due to diverse complex constraints and the vast number of possible combinations. To address such a problem, we introduce a novel formulation, combinatorial construction, which requires…

Cited by 22SourcePDFScholar
2021

Rebooting ACGAN: Auxiliary Classifier GANs with Stable Training

NeurIPS 2021poster

Conditional Generative Adversarial Networks (cGAN) generate realistic images by incorporating class information into GAN. While one of the most popular cGANs is an auxiliary classifier GAN with softmax cross-entropy loss (ACGAN), it is widely known that training ACGAN is challenging as the number of…

2021

Self-Calibrating Neural Radiance Fields

ICCV 2021poster

In this work, we propose a camera self-calibration algorithm for generic cameras with arbitrary non-linear distortions. We jointly learn the geometry of the scene and the accurate camera parameters without any calibration objects. Our camera model consists of a pinhole model, a fourth order radial d…

Cited by 268PDFcodeScholar
2020

HUMBI: A Large Multiview Dataset of Human Body Expressions

CVPR 2020poster

This paper presents a new large multiview dataset called HUMBI for human body expressions with natural clothing. The goal of HUMBI is to facilitate modeling view-specific appearance and geometry of gaze, face, hand, body, and garment from assorted people. 107 synchronized HD cam- eras are used to ca…

Cited by 111PDFScholar
2020

High-Dimensional Convolutional Networks for Geometric Pattern Recognition

CVPR 2020oral

High-dimensional geometric patterns appear in many computer vision problems. In this work, we present high-dimensional convolutional networks for geometric pattern recognition problems that arise in 2D and 3D registration problems. We first propose high-dimensional convolutional networks from 4 to 3…

Cited by 47PDFcodeScholar
2018

Tangent Convolutions for Dense Prediction in 3D

CVPR 2018poster

We present an approach to semantic scene analysis using deep convolutional networks. Our approach is based on tangent convolutions - a new construction for convolutional networks on 3D data. In contrast to volumetric approaches, our method operates directly on surface geometry. Crucially, the constr…

2016

Efficient and Robust Color Consistency for Community Photo Collections

CVPR 2016poster

We present an efficient technique to optimize color consistency of a collection of images depicting a common scene. Our method first recovers sparse pixel correspondences in the input images and stacks them into a matrix with many missing entries. We show that this matrix satisfies a rank two constr…

Cited by 70PDFScholar
2016

High-Quality Depth From Uncalibrated Small Motion Clip

CVPR 2016oral

We propose a novel approach that generates a high-quality depth map from a set of images captured with a small viewpoint variation, namely small motion clip. As opposed to prior methods that recover scene geometry and camera motions using pre-calibrated cameras, we introduce a self-calibrating bundl…

Cited by 132PDFcodeScholar
2016

Vision system and depth processing for DRC-HUBO+

ICRA 2016

This paper presents a vision system and a depth processing algorithm for DRC-HUBO+, the winner of the DRC finals 2015. Our system is designed to reliably capture 3D information of a scene and objects and to be robust to challenging environment conditions. We also propose a depth-map upsampling metho

Cited by 13SourceScholar
2015

Accurate Depth Map Estimation From a Lenslet Light Field Camera

CVPR 2015poster

This paper introduces an algorithm that accurately estimates depth maps using a lenslet light field camera. The proposed algorithm estimates the multi-view stereo correspondences with sub-pixel accuracy using the cost volume. The foundation for constructing accurate costs is threefold. First, the su…

Cited by 622SourcePDFScholar
2015

Multispectral Pedestrian Detection: Benchmark Dataset and Baseline

CVPR 2015poster

With the increasing interest in pedestrian detection, pedestrian datasets have also been the subject of research in the past decades. However, most existing datasets focus on a color channel, while a thermal channel is helpful for detection even in a dark environment. With this in mind, we propose a…

Cited by 1247SourcePDFScholar