← Search

Kyungdon Joo

33 accepted papers

2026

From Corners to Fiducial Tags: Revisiting Checkerboard Calibration for Event Cameras

CVPR 2026

The conventional checkerboard-based calibration for standard cameras faces fundamental limitations when applied to bio-inspired event cameras. Specifically, this stems from two challenges: (i) Events are triggered asynchronously at different timestamps along motion trajectories. If we accumulate the

Cited by 0SourcecodeScholar
2026

LightSplat: Fast and Memory-Efficient Open-Vocabulary 3D Scene Understanding in Five Seconds

CVPR 2026

Open-vocabulary 3D scene understanding enables users to segment novel objects in complex 3D environments through natural language. However, existing approaches remain slow, memory-intensive, and overly complex due to iterative optimization and dense per-Gaussian feature assignments. To address this,

Cited by 0SourcecodeScholar
2026

Prior-Constrained Explorative Guidance for Generalization in Diffusion Motion Planning

ICRA 2026poster

Diffusion-based planners have achieved generalization comparable to classical planners by leveraging inference-time optimization through guidance. However, their limited ability to capture environmental variations often constrains their responsiveness in unseen settings. In addition, the diversity-c…

Cited by 0codeScholar
2025

HUSH: Holistic Panoramic 3D Scene Understanding using Spherical Harmonics

CVPR 2025poster

Motivated by the efficiency of spherical harmonics (SH) in representing various physical phenomena, we propose a Holistic panoramic 3D scene Understanding framework using Spherical Harmonics, dubbed as HUSH. Our approach focuses on a unified framework adaptable to various 3D scene understanding task…

Cited by 0SourcePDFScholar
2025

San Francisco World: Leveraging Structural Regularities of Slope for 3-DoF Visual Compass

RA-L 2025

We propose the San Francisco world (SFW) model, a novel structural model inspired by San Francisco's hilly terrain, enabling 3D inter-floor navigation in urban areas rather than being limited to 2D intra-floor navigation of various robotics platforms. Our SFW consists of a single vertical dominant d

Cited by 4SourcecodeScholar
2025

VPOcc: Exploiting Vanishing Point for 3D Semantic Occupancy Prediction

IROS 2025

Understanding 3D scenes semantically and spatially is crucial for the safe navigation of robots and autonomous vehicles, aiding obstacle avoidance and accurate trajectory planning. Camera-based 3D semantic occupancy prediction, which infers complete voxel grids from 2D images, is gaining importance

Cited by 1SourcecodeScholar
2024

A Benchmark Dataset for Collaborative SLAM in Service Environments

RA-L 2024

We introduce a new multi-modal collaborative SLAM (C-SLAM) dataset for multiple service robots in various indoor service environments, called <monospace xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">C</monospace>-SLAM dataset in <monospace xmlns:mml="http:

Cited by 4SourcecodeScholar
2024

ContactGen: Contact-Guided Interactive 3D Human Generation for Partners

AAAI 2024technical

Among various interactions between humans, such as eye contact and gestures, physical interactions by contact can act as an essential moment in understanding human behaviors. Inspired by this fact, given a 3D partner human with the desired interaction label, we introduce a new task of 3D human gener…

Cited by 2SourcePDFScholar
2023

Diffusion-Based Signed Distance Fields for 3D Shape Generation

CVPR 2023poster

We propose a 3D shape generation framework (SDF-Diffusion in short) that uses denoising diffusion models with continuous 3D representation via signed distance fields (SDF). Unlike most existing methods that depend on discontinuous forms, such as point clouds, SDF-Diffusion generates high-resolution…

2023

Pose-Guided 3D Human Generation in Indoor Scene

AAAI 2023technical

In this work, we address the problem of scene-aware 3D human avatar generation based on human-scene interactions. In particular, we pay attention to the fact that physical contact between a 3D human and a scene (i.e., physical human-scene interactions) requires a geometrical alignment to generate na…

2023

SlaBins: Fisheye Depth Estimation using Slanted Bins on Road Environments

ICCV 2023poster

Although 3D perception for autonomous vehicles has focused on frontal-view information, more than half of fatal accidents occur due to side impacts in practice (e.g., T-bone crash). Motivated by this fact, we investigate the problem of side-view depth estimation, especially for monocular fisheye cam…

Cited by 6PDFScholar
2022

Adaptive Cost Volume Fusion Network for Multi-Modal Depth Estimation in Changing Environments

RA-L 2022

In this letter, we propose an adaptive cost volume fusion algorithm for multi-modal depth estimation in changing environments. Our method takes measurements from multi-modal sensors to exploit their complementary characteristics and generates depth cues from each modality in the form of adaptive cos

Cited by 13SourceScholar
2022

Quasi-Globally Optimal and Real-Time Visual Compass in Manhattan Structured Environments

RA-L 2022

We present a drift-free visual compass for estimating the three degrees of freedom (DoF) rotational motion of a camera by recognizing structural regularities in a Manhattan world (MW), which posits that the major structures conform to three orthogonal principal directions. Existing Manhattan frame e

Cited by 12SourcecodeScholar
2021

Learning Icosahedral Spherical Probability Map Based on Bingham Mixture Model for Vanishing Point Estimation

ICCV 2021poster

Existing vanishing point (VP) estimation methods rely on pre-extracted image lines and/or prior knowledge of the number of VPs. However, in practice, this information may be insufficient or unavailable. To solve this problem, we propose a network that treats a perspective image as input and predicts…

Cited by 9PDFScholar
2021

Volumetric Propagation Network: Stereo-LiDAR Fusion for Long-Range Depth Estimation

RA-L 2021

Stereo-LiDAR fusion is a promising task in that we can utilize two different types of 3D perceptions for practical usage - dense 3D information (stereo cameras) and highly-accurate sparse point clouds (LiDAR). However, due to their different modalities and structures, the method of aligning sensor d

Cited by 49SourceScholar
2020

Globally Optimal Relative Pose Estimation for Camera on a Selfie Stick

ICRA 2020poster

Taking selfies has become a photographic trend nowadays. We envision the emergence of the "video selfie" capturing a short continuous video clip (or burst photography) of the user, themselves. A selfie stick is usually used, whereby a camera is mounted on a stick for taking selfie photos. In this sc…

Cited by 2SourceScholar
2020

Globally Optimal and Efficient Vanishing Point Estimation in Atlanta World

ECCV 2020poster

Atlanta world holds for the scenes composed of a vertical dominant direction and several horizontal dominant directions. Vanishing point (VP) is the intersection of the image lines projected from parallel 3D lines. In Atlanta world, given a set of image lines, we aim to cluster them by the unknown-b…

Cited by 17SourcePDFScholar
2020

Non-Local Spatial Propagation Network for Depth Completion

ECCV 2020poster

In this paper, we propose a robust and efficient end-to-end non-local spatial propagation network for depth completion. The proposed network takes RGB and sparse depth images as inputs and estimates non-local neighbors and their affinities of each pixel, as well as an initial depth map with pixel-wi…

2020

SideGuide:A Large-scale Sidewalk Dataset for Guiding Impaired People

IROS 2020poster

In this paper, we introduce a new large-scale sidewalk dataset called SideGuide that could potentially help impaired people. Unlike most previous datasets, which are focused on road environments, we paid attention to sidewalks, where understanding the environment could provide the potential for impr…

Cited by 23SourceScholar
2019

Segment2Regress: Monocular 3D Vehicle Localization in Two Stages

RSS 2019poster

High-quality depth information is required to perform 3D vehicle detection, consequently, there exists a large performance gap between camera and LiDAR-based approaches. In this paper, our monocular camera-based 3D vehicle localization method alleviates the dependency on high-quality depth maps by t…

2019

Vehicular Multi-Camera Sensor System for Automated Visual Inspection of Electric Power Distribution Equipment

IROS 2019poster

In this paper, we present a multi-camera sensor system along with its control algorithm for automated visual inspection from a moving vehicle. To accomplish this task, we propose a unique hardware configuration consisting of a frontal stereo vision system, six lateral cameras motorized to tilt, and…

Cited by 7SourceScholar
2018

Globally Optimal Inlier Set Maximization for Atlanta Frame Estimation

CVPR 2018poster

In this work, we describe man-made structures via an appropriate structure assumption, called Atlanta world, which contains a vertical direction (typically the gravity direction) and a set of horizontal directions orthogonal to the vertical direction. Contrary to the commonly used Manhattan world as…

Cited by 22SourcePDFScholar
2017

Personalized Cinemagraphs Using Semantic Understanding and Collaborative Learning

ICCV 2017poster

Cinemagraphs are a compelling way to convey dynamic aspects of a scene. In these media, dynamic and still elements are juxtaposed to create an artistic and narrative experience. Creating a high-quality, aesthetically pleasing cinemagraph requires isolating objects in a semantically meaningful way an…

Cited by 19PDFScholar
2016

Vision system and depth processing for DRC-HUBO+

ICRA 2016

This paper presents a vision system and a depth processing algorithm for DRC-HUBO+, the winner of the DRC finals 2015. Our system is designed to reliably capture 3D information of a scene and objects and to be robust to challenging environment conditions. We also propose a depth-map upsampling metho

Cited by 13SourceScholar
2015

Accurate Camera Calibration Robust to Defocus Using a Smartphone

ICCV 2015poster

We propose a novel camera calibration method for defocused images using a smartphone under the assumption that the defocus blur is modeled as a convolution of a sharp image with a Gaussian point spread function (PSF). In contrast to existing calibration approaches which require well-focused images,…

Cited by 45PDFScholar
2015

High Quality Structure From Small Motion for Rolling Shutter Cameras

ICCV 2015poster

We present a practical 3D reconstruction method to obtain a high-quality dense depth map from narrow-baseline image sequences captured by commercial digital cameras, such as DSLRs or mobile phones. Depth estimation from small motion has gained interest as a means of various photographic editing, but…

Cited by 53PDFScholar