← Search

Sunghoon Im

38 accepted papers

2026

A Training-Free Style-Personalization via SVD-Based Feature Decomposition

CVPR 2026

We present a training-free framework for style-personalized image generation that operates during inference using a scale-wise autoregressive model. Our method generates a stylized image guided by a single reference style while preserving semantic consistency and mitigating content leakage. Through

Cited by 0SourceScholar
2026

Infinite-Story: A Training-Free Consistent Text-to-Image Generation

AAAI 2026technical

We present Infinite-Story, a training-free framework for consistent text-to-image (T2I) generation tailored for multi-prompt storytelling scenarios. Built upon a scale-wise autoregressive model, our method addresses two key challenges in consistent T2I generation: identity inconsistency and style in

Cited by 0SourcePDFScholar
2026

Scale-Invariant and View-Relational Representation Learning for Full Surround Monocular Depth

RA-L 2026

Recent foundation models demonstrate strong generalization capabilities in monocular depth estimation. However, directly applying these models to Full Surround Monocular Depth Estimation (FSMDE) presents two major challenges: (1) high computational cost, which limits real-time performance, and (2) d

Cited by 0SourceScholar
2026

Scale-Invariant and View-Relational Representation Learning for Full Surround Monocular Depth

ICRA 2026poster

Recent foundation models demonstrate strong generalization capabilities in monocular depth estimation. However, directly applying these models to Full Surround Monocular Depth Estimation (FSMDE) presents two major challenges: (1) high computational cost, which limits real-time performance, and (2) d…

2026

TaskForce: Cooperative Multi-agent Reinforcement Learning for Multi-task Optimization

CVPR 2026

Multi-task learning (MTL) involves the simultaneous optimization of multiple task-specific losses, often leading to gradient conflicts and scale imbalances that result in negative transfer. While existing multi-task optimization methods attempt to mitigate these challenges, they either lack the stoc

Cited by 0SourceScholar
2025

Flow4D: Leveraging 4D Voxel Network for LiDAR Scene Flow Estimation

RA-L 2025

Understanding the motion states of the surrounding environment is critical for safe autonomous driving. These motion states can be accurately derived from scene flow, which captures the three-dimensional motion field of points. Existing LiDAR scene flow methods extract spatial features from each poi

Cited by 20SourcecodeScholar
2025

Intrinsic Image Decomposition for Robust Self-supervised Monocular Depth Estimation on Reflective Surfaces

AAAI 2025technical

Self-supervised monocular depth estimation (SSMDE) has gained attention in the field of deep learning as it estimates depth without requiring ground truth depth maps. This approach typically uses a photometric consistency loss between a synthesized image, generated from the estimated depth, and the…

Cited by 0SourcePDFScholar
2025

JPEG Processing Neural Operator for Backward-Compatible Coding

ICCV 2025poster

Despite significant advances in learning-based lossy compression algorithms, standardizing codecs remains a critical challenge. In this paper, we present the JPEG Processing Neural Operator (JPNeO), a next-generation JPEG algorithm that maintains full backward compatibility with the current JPEG for…

2025

LOMM: Latest Object Memory Management for Temporally Consistent Video Instance Segmentation

ICCV 2025poster

In this paper, we present Latest Object Memory Management (LOMM) for temporally consistent video instance segmentation that significantly improves long-term instance tracking. At the core of our method is Latest Object Memory (LOM), which robustly tracks and continuously updates the latest states of…

Cited by 0SourcePDFScholar
2025

Self-supervised Monocular Depth Estimation Robust to Reflective Surface Leveraged by Triplet Mining

ICLR 2025poster

Self-supervised monocular depth estimation (SSMDE) aims to predict the dense depth map of a monocular image, by learning depth from RGB image sequences, eliminating the need for ground-truth depth labels. Although this approach simplifies data acquisition compared to supervised methods, it struggles…

Cited by 1SourcePDFScholar
2025

Style-Editor: Text-driven Object-centric Style Editing

CVPR 2025highlight

We present Text-driven object-centric style editing model named Style-Editor, a novel method that guides style editing at an object-centric level using textual inputs.The core of Style-Editor is our Patch-wise Co-Directional (PCD) loss, meticulously designed for precise object-centric editing that a…

Cited by 0SourcePDFScholar
2025

Towards Lossless Implicit Neural Representation via Bit Plane Decomposition

CVPR 2025poster

We quantify the upper bound on the size of the implicit neural representation (INR) model from a digital perspective. The upper bound of the model size increases exponentially as the required bit-precision increases. To this end, we present a bit-plane decomposition method that makes INR predict bit…

2024

BurstM: Deep Burst Multi-scale SR using Fourier Space with Optical Flow

ECCV 2024poster

"Multi frame super-resolution (MFSR) achieves higher performance than single image super-resolution (SISR), because MFSR leverages abundant information from multiple frames. Recent MFSR approaches adapt the deformable convolution network (DCN) to align the frames. However, the existing MFSR suffers…

2024

Density-aware Domain Generalization for LiDAR Semantic Segmentation

IROS 2024poster

3D LiDAR-based perception has made remarkable advancements, leading to the widespread adoption of LiDAR in autonomous driving systems. Despite these technological strides, variations in LiDAR sensors and environmental conditions can significantly deteriorate the performance of perception models, pri…

Cited by 2SourceScholar
2024

JDEC: JPEG Decoding via Enhanced Continuous Cosine Coefficients

CVPR 2024poster

We propose a practical approach to JPEG image decoding utilizing a local implicit neural representation with continuous cosine formulation. The JPEG algorithm significantly quantizes discrete cosine transform (DCT) spectra to achieve a high compression rate inevitably resulting in quality degradatio…

2024

Multi-task Learning for Real-time Autonomous Driving Leveraging Task-adaptive Attention Generator

ICRA 2024poster

Real-time processing is crucial in autonomous driving systems due to the imperative of instantaneous decision-making and rapid response. In real-world scenarios, autonomous vehicles are continuously tasked with interpreting their surroundings, analyzing intricate sensor data, and making decisions wi…

Cited by 3SourceScholar
2023

Deep Digging into the Generalization of Self-Supervised Monocular Depth Estimation

AAAI 2023technical

Self-supervised monocular depth estimation has been widely studied recently. Most of the work has focused on improving performance on benchmark datasets, such as KITTI, but has offered a few experiments on generalization performance. In this paper, we investigate the backbone networks (e.g., CNNs, T…

2023

Depth-discriminative Metric Learning for Monocular 3D Object Detection

NeurIPS 2023poster

Monocular 3D object detection poses a significant challenge due to the lack of depth information in RGB images. Many existing methods strive to enhance the object depth estimation performance by allocating additional parameters for object depth estimation, utilizing extra modules or data. In contras…

Cited by 5SourcePDFScholar
2023

Dynamic Neural Network for Multi-Task Learning Searching Across Diverse Network Topologies

CVPR 2023poster

In this paper, we present a new MTL framework that searches for structures optimized for multiple tasks with diverse graph topologies and shares features among tasks. We design a restricted DAG-based central network with read-in/read-out layers to build topologically diverse task-adaptive structures…

Cited by 8SourcePDFScholar
2022

ADAS: A Direct Adaptation Strategy for Multi-Target Domain Adaptive Semantic Segmentation

CVPR 2022poster

In this paper, we present a direct adaptation strategy (ADAS), which aims to directly adapt a single model to multiple target domains in a semantic segmentation task without pretrained domain-specific models. To do so, we design a multi-target domain transfer network (MTDT-Net) that aligns visual at…

Cited by 29PDFcodeScholar
2022

Facial Depth and Normal Estimation Using Single Dual-Pixel Camera

ECCV 2022poster

"Recently, Dual-Pixel (DP) sensors have been adopted in many imaging devices. However, despite their various advantages, DP sensors are used just for faster auto-focus and aesthetic image captures, and research on their usage for 3D facial understanding has been limited due to the lack of datasets a…

2021

DRANet: Disentangling Representation and Adaptation Networks for Unsupervised Cross-Domain Adaptation

CVPR 2021poster

In this paper, we present DRANet, a network architecture that disentangles image representations and transfers the visual attributes in a latent space for unsupervised cross-domain adaptation. Unlike the existing domain adaptation methods that learn associated features sharing a domain, DRANet prese…

Cited by 85PDFcodeScholar
2021

Learning Monocular Depth in Dynamic Scenes via Instance-Aware Projection Consistency

AAAI 2021technical

We present an end-to-end joint training framework that explicitly models 6-DoF motion of multiple dynamic objects, ego-motion, and depth in a monocular camera setup without supervision. Our technical contributions are three-fold. First, we highlight the fundamental difference between inverse and for…

2021

VolumeFusion: Deep Depth Fusion for 3D Scene Reconstruction

ICCV 2021poster

To reconstruct a 3D scene from a set of calibrated views, traditional multi-view stereo techniques rely on two distinct stages: local depth maps computation and global depth maps fusion. Recent studies concentrate on deep neural architectures for depth estimation by using conventional depth fusion m…

Cited by 63PDFScholar
2020

Learning Shape-based Representation for Visual Localization in Extremely Changing Conditions

ICRA 2020poster

Visual localization is an important task for applications such as navigation and augmented reality, but is a challenging problem when there are changes in scene appearances through day, seasons, or environments. In this paper, we present a convolutional neural network (CNN)-based approach for visual…

Cited by 6SourceScholar
2019

DISC: A Large-scale Virtual Dataset for Simulating Disaster Scenarios

IROS 2019poster

In this paper, we present the first large-scale synthetic dataset for visual perception in disaster scenarios, and analyze state-of-the-art methods for multiple computer vision tasks with reference baselines. We simulated before and after disaster scenarios such as fire and building collapse for fif…

Cited by 15SourceScholar
2018

RANUS: RGB and NIR Urban Scene Dataset for Deep Scene Parsing

RA-L 2018

In this letter, we present a data-driven method for scene parsing of road scenes to utilize single-channel near-infrared (NIR) images. To overcome the lack of data problem in non-RGB spectrum, we define a new color space and decompose the task of deep scene parsing into two subtasks with two separat

Cited by 42SourceScholar
2017

Noise Robust Depth From Focus Using a Ring Difference Filter

CVPR 2017spotlight

Depth from focus (DfF) is a method of estimating depth of a scene by using the information acquired through the change of the focus of a camera. Within the framework of DfF, the focus measure (FM) forms the foundation on which the accuracy of the output is determined. With the result from the FM, th…

Cited by 47PDFScholar
2016

High-Quality Depth From Uncalibrated Small Motion Clip

CVPR 2016oral

We propose a novel approach that generates a high-quality depth map from a set of images captured with a small viewpoint variation, namely small motion clip. As opposed to prior methods that recover scene geometry and camera motions using pre-calibrated cameras, we introduce a self-calibrating bundl…

Cited by 132PDFcodeScholar
2016

Stereo Matching With Color and Monochrome Cameras in Low-Light Conditions

CVPR 2016poster

Consumer devices with stereo cameras have become popular because of their low-cost depth sensing capability. However, those systems usually suffer from low imaging quality and inaccurate depth acquisition under low-light conditions. To address the problem, we present a new stereo matching method wit…

Cited by 60PDFScholar
2015

High Quality Structure From Small Motion for Rolling Shutter Cameras

ICCV 2015poster

We present a practical 3D reconstruction method to obtain a high-quality dense depth map from narrow-baseline image sequences captured by commercial digital cameras, such as DSLRs or mobile phones. Depth estimation from small motion has gained interest as a means of various photographic editing, but…

Cited by 53PDFScholar