← Search

Gim Hee Lee

139 accepted papers

2026

Active3D: Active High-Fidelity 3D Reconstruction via Multi-Level Uncertainty Quantification

AAAI 2026technical

In this paper, we present an active exploration framework for high-fidelity 3D reconstruction that incrementally builds a multi-level uncertainty space and selects next-best-views through an uncertainty-driven motion planner. We introduce a hybrid implicit–explicit representation that fuses neural

Cited by 0SourcePDFScholar
2026

Condensed Data Expansion Using Model Inversion for Knowledge Distillation

AAAI 2026technical

Condensed datasets offer a compact representation of larger datasets, but training models directly on them or using them to enhance model performance through knowledge distillation (KD) can result in suboptimal outcomes due to limited information. To address this, we propose a method that expands co

Cited by 0SourcePDFScholar
2026

D3D-VLP: Dynamic 3D Vision-Language-Planning Model for Embodied Grounding and Navigation

CVPR 2026

Embodied agents face a critical dilemma that end-to-end models lack interpretability and explicit 3D reasoning, while modular systems ignore cross-component interdependencies and synergies. To bridge this gap, we propose the Dynamic 3D Vision-Language-Planning Model (D3D-VLP). Our model introduces t

Cited by 0SourcecodeScholar
2026

EmbodiedSplat: Online Feed-Forward Semantic 3DGS for Open-Vocabulary 3D Scene Understanding

CVPR 2026

Understanding a 3D scene immediately with its exploration is essential for embodied tasks, where an agent must construct and comprehend the 3D representation in an online and nearly real-time manner. In this study, we propose EmbodiedSplat, an online feed-forward 3DGS for open-vocabulary scene under

Cited by 0SourcecodeScholar
2026

From Obstacles to Etiquette: Robot Social Navigation With VLM-Informed Path Selection

RA-L 2026

Navigating socially in human environments requires more than satisfying geometric constraints, as collision-free paths may still interfere with ongoing activities or conflict with social norms. Addressing this challenge calls for analyzing interactions between agents and incorporating common-sense r

Cited by 1SourcecodeScholar
2026

HandMCM: Multi-modal Point Cloud-based Correspondence State Space Model for 3D Hand Pose Estimation

AAAI 2026technical

3D hand pose estimation that involves accurate estimation of 3D human hand keypoint locations is crucial for many human-computer interaction applications such as augmented reality. However, this task poses significant challenges due to self-occlusion of the hands and occlusions caused by interaction

Cited by 0SourcePDFScholar
2026

MotionScale: Reconstructing Appearance, Geometry, and Motion of Dynamic Scenes with Scalable 4D Gaussian Splatting

CVPR 2026

Realistic reconstruction of dynamic 4D scenes from monocular videos is essential for understanding the physical world. Despite recent progress in neural rendering, existing methods often struggle to recover accurate 3D geometry and temporally consistent motion in complex environments. To address the

Cited by 0SourcecodeScholar
2026

Revisiting Photometric Ambiguity for Accurate Gaussian-Splatting Surface Reconstruction

ICML 2026poster

Surface reconstruction with differentiable rendering has achieved impressive performance in recent years, yet the pervasive photometric ambiguities have strictly bottlenecked existing approaches. This paper presents AmbiSuR, a framework that explores an intrinsic solution upon Gaussian Splatting for…

Cited by 0SourceScholar
2026

RiemanLine: Riemannian Manifold Representation of 3D Lines for Factor Graph Optimization

AAAI 2026technical

Minimal parametrization of 3D lines plays a critical role in camera localization and structural mapping. Existing representations in robotics and computer vision predominantly handle independent lines, overlooking structural regularities such as sets of parallel lines that are pervasive in man-made

Cited by 0SourcePDFScholar
2026

Segment Any Events with Language

ICLR 2026poster

Scene understanding with free-form language has been widely explored within diverse modalities such as images, point clouds, and LiDAR. However, related studies on event sensors are scarce or narrowly centered on semantic-level understanding. We introduce **SEAL**, the first Semantic-aware Segment A…

Cited by 0SourceScholar
2026

TeamHOI: Learning a Unified Policy for Cooperative Human-Object Interactions with Any Team Size

CVPR 2026

Physics-based humanoid control has achieved remarkable progress in enabling realistic and high-performing single-agent behaviors, yet extending these capabilities to cooperative human-object interaction (HOI) remains challenging. We present TeamHOI, a framework that enables a single decentralized po

Cited by 0SourcecodeScholar
2025

4D3R: Motion-Aware Neural Reconstruction and Rendering of Dynamic Scenes from Monocular Videos

NeurIPS 2025poster

Novel view synthesis from monocular videos of dynamic scenes with unknown camera poses remains a fundamental challenge in computer vision and graphics. While recent advances in 3D representations such as Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have shown promising results for…

Cited by 0SourceScholar
2025

ComPC: Completing a 3D Point Cloud with 2D Diffusion Priors

ICLR 2025poster

3D point clouds directly collected from objects through sensors are often incomplete due to self-occlusion. Conventional methods for completing these partial point clouds rely on manually organized training sets and are usually limited to object categories seen during training. In this work, we prop…

2025

Deep Gaussian from Motion: Exploring 3D Geometric Foundation Models for Gaussian Splatting

NeurIPS 2025poster

Neural radiance fields (NeRF) and 3D Gaussian Splatting (3DGS) are popular techniques to reconstruct and render photorealistic images. However, the prerequisite of running Structure-from-Motion (SfM) to get camera poses limits their completeness. Although previous methods can reconstruct a few unpos…

Cited by 0SourceScholar
2025

DiET-GS: Diffusion Prior and Event Stream-Assisted Motion Deblurring 3D Gaussian Splatting

CVPR 2025poster

Reconstructing sharp 3D representations from blurry multi-view images are long-standing problem in computer vision. Recent works attempt to enhance high-quality novel view synthesis from the motion blur by leveraging event-based cameras, benefiting from high dynamic range and microsecond temporal re…

Cited by 0SourcePDFScholar
2025

DuCos: Duality Constrained Depth Super-Resolution via Foundation Model

ICCV 2025poster

We introduce DuCos, a novel depth super-resolution framework grounded in Lagrangian duality theory, offering a flexible integration of multiple constraints and reconstruction objectives to enhance accuracy and robustness. Our DuCos is the first to significantly improve generalization across diverse…

2025

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation

NeurIPS 2025oral

Vision-and-Language Navigation (VLN) is a core task where embodied agents leverage their spatial mobility to navigate in 3D environments toward designated destinations based on natural language instructions. Recently, video-language large models (Video-VLMs) with strong generalization capabilities a…

Cited by 0SourcecodeScholar
2025

Event-aided Dense and Continuous Point Tracking: Everywhere and Anytime

ICCV 2025poster

Recent point tracking methods have made great strides in recovering the trajectories of any point (especially key points) in long video sequences associated with large motions. However, the spatial and temporal granularities of point trajectories remain constrained by limited motion estimation accur…

Cited by 0SourcePDFScholar
2025

FlexEvent: Towards Flexible Event-Frame Object Detection at Varying Operational Frequencies

NeurIPS 2025poster

Event cameras offer unparalleled advantages for real-time perception in dynamic environments, thanks to the microsecond-level temporal resolution and asynchronous operation. Existing event detectors, however, are limited by fixed-frequency paradigms and fail to fully exploit the high-temporal resolu…

Cited by 0SourceScholar
2025

GEAL: Generalizable 3D Affordance Learning with Cross-Modal Consistency

CVPR 2025poster

Identifying affordance regions on 3D objects from semantic cues is essential for robotics and human-machine interaction. However, existing 3D affordance learning methods struggle with generalization and robustness due to limited annotated data and a reliance on 3D backbones focused on geometric enco…

2025

GenXD: Generating Any 3D and 4D Scenes

ICLR 2025poster

Recent developments in 2D visual generation have been remarkably successful. However, 3D and 4D generation remain challenging in real-world applications due to the lack of large-scale 4D data and effective model design. In this paper, we propose to jointly investigate general 3D and 4D generation by…

Cited by 8SourcePDFScholar
2025

Generalizable Human Gaussians from Single-View Image

ICLR 2025poster

In this work, we tackle the task of learning 3D human Gaussians from a single image, focusing on recovering detailed appearance and geometry including unobserved regions. We introduce a single-view generalizable Human Gaussian Model (HGM), which employs a novel generate-then-refine pipeline with the…

2025

LLaFEA: Frame-Event Complementary Fusion for Fine-Grained Spatiotemporal Understanding in LMMs

ICCV 2025poster

Large multimodal models (LMMs) excel in scene understanding but struggle with fine-grained spatiotemporal reasoning due to weak alignment between linguistic and visual representations. Existing methods map textual positions and durations into the visual space encoded from frame-based videos, but suf…

2025

Learnable Infinite Taylor Gaussian for Dynamic View Rendering

CVPR 2025poster

Capturing the temporal evolution of Gaussian properties such as position, rotation, and scale is a challenging task due to the vast number of time-varying parameters and the limited photometric data available, which generally results in convergence issues, making it difficult to find an optimal solu…

Cited by 0SourcePDFScholar
2025

LiLoc: Lifelong Localization Using Adaptive Submap Joining and Egocentric Factor Graph

ICRA 2025

This paper proposes a versatile graph-based lifelong localization framework using LiDAR, LiLoc, which enhances its timeliness by maintaining a single central session while improves the accuracy through multi-modal factors between the central and subsidiary sessions. First, an adaptive submap joining

Cited by 3SourcecodeScholar
2025

MAR-3D: Progressive Masked Auto-regressor for High-Resolution 3D Generation

CVPR 2025highlight

Recent advances in auto-regressive transformers have revolutionized generative modeling across different domains, from language processing to visual generation, demonstrating remarkable capabilities. However, applying these advances to 3D generation presents three key challenges: the unordered natur…

Cited by 1SourcePDFScholar
2025

NavRAG: Generating User Demand Instructions for Embodied Navigation through Retrieval-Augmented LLM

ACL 2025finding

Vision-and-Language Navigation (VLN) is an essential skill for embodied agents, allowing them to navigate in 3D environments following natural language instructions. High-performance navigation models require a large amount of training data, the high cost of manually annotating data has seriously hi…

2025

Neuralized Markov Random Field for Interaction-Aware Stochastic Human Trajectory Prediction

ICLR 2025poster

Interactive human motions and the continuously changing nature of intentions pose significant challenges for human trajectory prediction. In this paper, we present a neuralized Markov random field (MRF)-based motion evolution method for probabilistic interaction-aware human trajectory prediction. We…

2025

PoseCrafter: Extreme Pose Estimation with Hybrid Video Synthesis

NeurIPS 2025poster

Pairwise camera pose estimation from sparsely overlapping image pairs remains a critical and unsolved challenge in 3D vision. Most existing methods struggle with image pairs that have small or no overlap. Recent approaches attempt to address this by synthesizing intermediate frames using video inte…

Cited by 0SourceScholar
2025

Reconstructing Close Human Interaction with Appearance and Proxemics Reasoning

CVPR 2025poster

Due to visual ambiguities and inter-person occlusions, existing human pose estimation methods cannot recover plausible close interactions from in-the-wild videos. Even state-of-the-art large foundation models (e.g., SAM) cannot accurately distinguish human semantics in such challenging scenarios. In…

Cited by 0SourcePDFScholar
2025

STD-GS: Exploring Frame-Event Interaction for SpatioTemporal-Disentangled Gaussian Splatting to Reconstruct High-Dynamic Scene

ICCV 2025poster

High-dynamic scene reconstruction aims to represent static background with rigid spatial features and dynamic objects with deformed continuous spatiotemporal features. Typically, existing methods adopt unified representation model (e.g., Gaussian) to directly match the spatiotemporal features of dyn…

2025

Statistical Confidence Rescoring for Robust 3D Scene Graph Generation from Multi-View Images

ICCV 2025poster

Modern 3D semantic scene graph estimation methods utilize ground truth 3D annotations to accurately predict target objects, predicates, and relationships. In the absence of given 3D ground truth representations, we explore leveraging only multi-view RGB images to tackle this task. To attain robust f…

2025

Street Gaussians without 3D Object Tracker

ICCV 2025poster

Realistic scene reconstruction in driving scenarios poses significant challenges due to fast-moving objects. Most existing methods rely on labor-intensive manual labeling of object poses to reconstruct dynamic objects in canonical space and move them based on these poses during rendering. While some…

Cited by 0SourcePDFScholar
2025

TreeMeshGPT: Artistic Mesh Generation with Autoregressive Tree Sequencing

CVPR 2025poster

We introduce TreeMeshGPT, an autoregressive Transformer designed to generate high-quality artistic meshes aligned with input point clouds. Instead of the conventional next-token prediction in autoregressive Transformer, we propose a novel Autoregressive Tree Sequencing where the next input token is…

2025

X-Scene: Large-Scale Driving Scene Generation with High Fidelity and Flexible Controllability

NeurIPS 2025poster

Diffusion models are advancing autonomous driving by enabling realistic data synthesis, predictive end-to-end planning, and closed-loop simulation, with a primary focus on temporally consistent generation. However, large-scale 3D scene generation requiring spatial coherence remains underexplored. In…

Cited by 0SourceScholar
2024

Closely Interactive Human Reconstruction with Proxemics and Physics-Guided Adaption

CVPR 2024poster

Existing multi-person human reconstruction approaches mainly focus on recovering accurate poses or avoiding penetration but overlook the modeling of close interactions. In this work we tackle the task of reconstructing closely interactive humans from a monocular video. The main challenge of this tas…

2024

DOGS: Distributed-Oriented Gaussian Splatting for Large-Scale 3D Reconstruction Via Gaussian Consensus

NeurIPS 2024poster

The recent advances in 3D Gaussian Splatting (3DGS) show promising results on the novel view synthesis (NVS) task. With its superior rendering performance and high-fidelity rendering quality, 3DGS is excelling at its previous NeRF counterparts. The most recent 3DGS method focuses either on improving…

2024

Deblur e-NeRF: NeRF from Motion-Blurred Events under High-speed or Low-light Conditions

ECCV 2024poster

"The distinctive design philosophy of event cameras makes them ideal for high-speed, high dynamic range & low-light environments, where standard cameras underperform. However, event cameras also suffer from motion blur, especially under these challenging conditions, contrary to what most think. This…

Cited by 6SourcePDFScholar
2024

ESCAPE: Encoding Super-keypoints for Category-Agnostic Pose Estimation

CVPR 2024poster

In this paper we tackle the task of category-agnostic pose estimation (CAPE) which aims to predict poses for objects of any category with few annotated samples. Previous works either rely on local matching between features of support and query samples or require support keypoint identifier. The form…

2024

FreeSplat: Generalizable 3D Gaussian Splatting Towards Free View Synthesis of Indoor Scenes

NeurIPS 2024poster

Empowering 3D Gaussian Splatting with generalization ability is appealing. However, existing generalizable 3D Gaussian Splatting methods are largely confined to narrow-range interpolation between stereo images due to their heavy backbones, thus lacking the ability to accurately localize 3D Gaussian…

2024

GeoGaussian: Geometry-aware Gaussian Splatting for Scene Rendering

ECCV 2024poster

"During the Gaussian Splatting optimization process, the scene geometry can gradually deteriorate if its structure is not deliberately preserved, especially in non-textured regions such as walls, ceilings, and furniture surfaces. This degradation significantly affects the rendering quality of novel…

Cited by 25SourcePDFScholar
2024

Learning to Decouple the Lights for 3D Face Texture Modeling

NeurIPS 2024poster

Existing research has made impressive strides in reconstructing human facial shapes and textures from images with well-illuminated faces and minimal external occlusions. Nevertheless, it remains challenging to recover accurate facial textures from scenarios with complicated illumination affected by…

Cited by 0SourcePDFScholar
2024

MVSDet: Multi-View Indoor 3D Object Detection via Efficient Plane Sweeps

NeurIPS 2024poster

The key challenge of multi-view indoor 3D object detection is to infer accurate geometry information from images for precise 3D detection. Previous method relies on NeRF for geometry reasoning. However, the geometry extracted from NeRF is generally inaccurate, which leads to sub-optimal detection pe…

2024

TreeSBA: Tree-Transformer for Self-Supervised Sequential Brick Assembly

ECCV 2024poster

"Inferring step-wise actions to assemble 3D objects with primitive bricks from images is a challenging task due to complex constraints and the vast number of possible combinations. Recent studies have demonstrated promising results on sequential LEGO brick assembly through the utilization of LEGO-Gr…

2024

UNIKD: UNcertainty-Filtered Incremental Knowledge Distillation for Neural Implicit Representation

ECCV 2024poster

"Recent neural implicit representations (NIRs) have achieved great success in the tasks of 3D reconstruction and novel view synthesis. However, they require the images of a scene from different camera views to be available for one-time training. This is expensive especially for scenarios with large-…

2024

URS-NeRF: Unordered Rolling Shutter Bundle Adjustment for Neural Radiance Fields

ECCV 2024poster

"In this paper, we propose a novel rolling shutter bundle adjustment method for neural radiance fields (NeRF), which utilizes the unordered rolling shutter (RS) images to obtain the implicit 3D representation. Existing NeRF methods suffer from low-quality images and inaccurate initial camera poses d…

Cited by 0SourcePDFScholar
2024

VCR-GauS: View Consistent Depth-Normal Regularizer for Gaussian Surface Reconstruction

NeurIPS 2024poster

Although 3D Gaussian Splatting has been widely studied because of its realistic and efficient novel-view synthesis, it is still challenging to extract a high-quality surface from the point-based representation. Previous works improve the surface by incorporating geometric priors from the off-the-she…

Cited by 14SourcePDFScholar
2023

AdaSfM: From Coarse Global to Fine Incremental Adaptive Structure from Motion

ICRA 2023poster

Despite the impressive results achieved by many existing Structure from Motion (SfM) approaches, there is still a need to improve the robustness, accuracy, and efficiency on large-scale scenes with many outlier matches and sparse view graphs. In this paper, we propose AdaSfM: a coarse-to-fine adapti…

Cited by 10SourceScholar
2023

Boosting Long-tailed Object Detection via Step-wise Learning on Smooth-tail Data

ICCV 2023poster

Real-world data tends to follow a long-tailed distribution, where the class imbalance results in dominance of the head classes during training. In this paper, we propose a frustratingly simple but effective step-wise learning framework to gradually enhance the capability of the model in detecting al…

Cited by 5PDFScholar
2023

ContraNeRF: Generalizable Neural Radiance Fields for Synthetic-to-Real Novel View Synthesis via Contrastive Learning

CVPR 2023poster

Although many recent works have investigated generalizable NeRF-based novel view synthesis for unseen scenes, they seldom consider the synthetic-to-real generalization, which is desired in many practical applications. In this work, we first investigate the effects of synthetic data in synthetic-to-r…

2023

Incremental-DETR: Incremental Few-Shot Object Detection via Self-Supervised Learning

AAAI 2023technical

Incremental few-shot object detection aims at detecting novel classes without forgetting knowledge of the base classes with only a few labeled training data from the novel classes. Most related prior works are on incremental object detection that rely on the availability of abundant training samples…

2023

NU-MCC: Multiview Compressive Coding with Neighborhood Decoder and Repulsive UDF

NeurIPS 2023poster

Remarkable progress has been made in 3D reconstruction from single-view RGB-D inputs. MCC is the current state-of-the-art method in this field, which achieves unprecedented success by combining vision Transformers with large-scale training. However, we identified two key limitations of MCC: 1) The T…

2023

Optimizing the Placement of Roadside LiDARs for Autonomous Driving

ICCV 2023poster

Multi-agent cooperative perception is an increasingly popular topic in the field of autonomous driving, where roadside LiDARs play an essential role. However, how to optimize the placement of roadside LiDARs is a crucial but often overlooked problem. This paper proposes an approach to optimize the p…

Cited by 16PDFScholar
2023

Overcoming the Trade-Off Between Accuracy and Plausibility in 3D Hand Shape Reconstruction

CVPR 2023poster

Direct mesh fitting for 3D hand shape reconstruction estimates highly accurate meshes. However, the resulting meshes are prone to artifacts and do not appear as plausible hand shapes. Conversely, parametric models like MANO ensure plausible hand shapes but are not as accurate as the non-parametric m…

Cited by 9SourcePDFScholar
2023

Unsupervised Feature Representation Learning for Domain-generalized Cross-domain Image Retrieval

ICCV 2023poster

Cross-domain image retrieval has been extensively studied due to its high practical value. In recently proposed unsupervised cross-domain image retrieval methods, efforts are taken to break the data annotation barrier. However, applicability of the model is still confined to domains seen during trai…

Cited by 8PDFcodeScholar
2022

Adversarial Style Augmentation for Domain Generalized Urban-Scene Segmentation

NeurIPS 2022accept

In this paper, we consider the problem of domain generalization in semantic segmentation, which aims to learn a robust model using only labeled synthetic (source) data. The model is expected to perform well on unseen real (target) domains. Our study finds that the image style variation can largely i…

Cited by 94SourcePDFScholar
2022

Minimal Neural Atlas: Parameterizing Complex Surfaces with Minimal Charts and Distortion

ECCV 2022poster

"Explicit neural surface representations allow for exact and efficient extraction of the encoded surface at arbitrary precision, as well as analytic derivation of differential geometric properties such as surface normal and curvature. Such desirable properties, which are absent in its implicit count…

2022

Rethinking IoU-Based Optimization for Single-Stage 3D Object Detection

ECCV 2022poster

"Since Intersection-over-Union (IoU) based optimization maintains the consistency of the final IoU prediction metric and losses, it has been widely used in both regression and classification branches of single-stage 2D object detectors. Recently, several 3D object detection methods adopt IoU-based o…

2022

Style-Hallucinated Dual Consistency Learning for Domain Generalized Semantic Segmentation

ECCV 2022poster

"In this paper, we study the task of synthetic-to-real domain generalized semantic segmentation, which aims to learn a model that is robust to unseen real-world scenes using only synthetic data. The large domain shift between synthetic and real-world data, including the limited source environmental…

2022

Teaching with Soft Label Smoothing for Mitigating Noisy Labels in Facial Expressions

ECCV 2022poster

"Recent studies have highlighted the problem of noisy labels in large scale in-the-wild facial expressions datasets due to the uncertainties caused by ambiguous facial expressions, low-quality facial images, and the subjectiveness of annotators. To solve the problem of noisy labels, we propose Soft…

2021

A General Framework for Lifelong Localization and Mapping in Changing Environment

IROS 2021poster

The environment of most real-world scenarios such as malls and supermarkets changes at all times. A pre-built map that does not account for these changes becomes out-of-date easily. Therefore, it is necessary to have an up-to-date model of the environment to facilitate long-term operation of a robot…

Cited by 46SourcecodeScholar
2021

Bridging Non Co-occurrence with Unlabeled In-the-wild Data for Incremental Object Detection

NeurIPS 2021poster

Deep networks have shown remarkable results in the task of object detection. However, their performance suffers critical drops when they are subsequently trained on novel classes without any sample from the base classes originally used to train the model. This phenomenon is known as catastrophic for…

2021

MINE: Towards Continuous Depth MPI With NeRF for Novel View Synthesis

ICCV 2021poster

In this paper, we propose MINE to perform novel view synthesis and depth estimation via dense 3D reconstruction from a single image. Our approach is a continuous depth generalization of the Multiplane Images (MPI) by introducing the NEural radiance fields (NeRF). Given a single image as input, MINE…

Cited by 171PDFcodeScholar
2020

Multi-person 3D Pose Estimation in Crowded Scenes Based on Multi-View Geometry

ECCV 2020poster

Epipolar constraints are at the core of feature matching and depth estimation in current multi-person multi-camera 3D human pose estimation methods. Despite the satisfactory performance of this formulation in sparser crowd scenes, its effectiveness is frequently challenged under denser crowd circums…

2020

PointAtrousGraph: Deep Hierarchical Encoder-Decoder with Point Atrous Convolution for Unorganized 3D Points

ICRA 2020poster

Motivated by the success of encoding multi-scale contextual information for image analysis, we propose our PointAtrousGraph (PAG) - a deep permutation-invariant hierarchical encoder-decoder for efficiently exploiting multi-scale edge features in point clouds. Our PAG is constructed by several novel…

Cited by 37SourcecodeScholar
2020

Relative Pose Estimation of Calibrated Cameras with Known SE(3) Invariants

ECCV 2020poster

The $\mathrm{SE}(3)$ invariants of a pose include its rotation angle and screw translation. In this paper, we present a complete comprehensive study of the relative pose estimation problem for a calibrated camera constrained by known $\mathrm{SE}(3)$ invariant, which involves 5 minimal problems in t…

2020

Shape Prior Deformation for Categorical 6D Object Pose and Size Estimation

ECCV 2020poster

We present a novel learning approach to recover the 6D poses and sizes of unseen object instances from an RGB-D image. To handle the intra-class shape variation, we propose a deep network to reconstruct the 3D object model by explicitly modeling the deformation from a pre-learned categorical shape p…

2019

2D3D-Matchnet: Learning To Match Keypoints Across 2D Image And 3D Point Cloud

ICRA 2019poster

Large-scale point cloud generated from 3D sensors is more accurate than its image-based counterpart. However, it is seldom used in visual pose estimation due to the difficulty in obtaining 2D-3D image to point cloud correspondences. In this paper, we propose the 2D3D-MatchNet - an end-to-end deep ne…

Cited by 146SourceScholar
2019

Degeneracy in Self-Calibration Revisited and a Deep Learning Solution for Uncalibrated SLAM

IROS 2019poster

Self-calibration of camera intrinsics and radial distortion has a long history of research in the computer vision community. However, it remains rare to see real applications of such techniques to modern Simultaneous Localization And Mapping (SLAM) systems, especially in driving scenarios. In this p…

Cited by 27SourceScholar
2019

Project AutoVision: Localization and 3D Scene Perception for an Autonomous Vehicle with a Multi-Camera System

ICRA 2019poster

Project AutoVision aims to develop localization and 3D scene perception capabilities for a self-driving vehicle. Such capabilities will enable autonomous navigation in urban and rural environments, in day and night, and with cameras as the only exteroceptive sensors. The sensor suite employs many ca…

Cited by 149SourceScholar
2018

CVM-Net: Cross-View Matching Network for Image-Based Ground-to-Aerial Geo-Localization

CVPR 2018poster

The problem of localization on a geo-referenced aerial/satellite map given a query ground view image remains challenging due to the drastic change in viewpoint that causes traditional image descriptors based matching to fail. We leverage on the recent success of deep learning to propose the CVM-Net…

2018

Convolutional Sequence to Sequence Model for Human Dynamics

CVPR 2018poster

Human motion modeling is a classic problem in com- puter vision and graphics. Challenges in modeling human motion include high dimensional prediction as well as extremely complicated dynamics.We present a novel approach to human motion modeling based on convolutional neural networks (CNN). The hiera…

2018

PointNetVLAD: Deep Point Cloud Based Retrieval for Large-Scale Place Recognition

CVPR 2018poster

Unlike its image based counterpart, point cloud based retrieval for place recognition has remained as an unexplored and unsolved problem. This is largely due to the difficulty in extracting local feature descriptors from a point cloud that can subsequently be encoded into a global descriptor for th…

2016

Object detection and motion planning for automated welding of tubular joints

IROS 2016poster

Automatic welding of tubular TKY joints is an important and challenging task for the marine and offshore industry. In this paper, a framework for tubular joint detection and motion planning is proposed. The pose of the real tubular joint is detected using RGB-D sensors, which is used to obtain a rea…

Cited by 27SourceScholar
2015

Line-Sweep: Cross-Ratio For Wide-Baseline Matching and 3D Reconstruction

CVPR 2015poster

We propose a simple and useful idea based on cross-ratio constraint for wide-baseline matching and 3D reconstruction. Most existing methods exploit feature points and planes from images. Lines have always been considered notorious for both matching and reconstruction due to the lack of good line des…

Cited by 50SourcePDFScholar