← Search

Hongdong Li

114 accepted papers

2026

BiFM: Bidirectional Flow Matching for Few-Step Image Editing and Generation

CVPR 2026

Recent diffusion and flow matching models have demonstrated strong capabilities in image generation and editing by progressively removing noise through iterative sampling. While this enables flexible inversion for semantic-preserving edits, few-step sampling regimes suffer from poor forward process

Cited by 1SourceScholar
2026

DiffNR: Diffusion-Enhanced Neural Representation Optimization for Sparse-View 3D Tomographic Reconstruction

AAAI 2026technical

Neural representations (NRs), such as neural fields and 3D Gaussians, effectively model volumetric data in computed tomography (CT) but suffer from severe artifacts under sparse-view settings. To address this, we propose DiffNR, a novel framework that enhances NR optimization with diffusion priors.

Cited by 0SourcePDFScholar
2026

Joint Shadow Generation and Relighting via Light-Geometry Interaction Maps

ICLR 2026poster

We propose Light–Geometry Interaction (LGI) maps, a novel representation that encodes light-aware occlusion from monocular depth. Unlike ray tracing, which requires full 3D reconstruction, LGI captures essential light–shadow interactions reliably and accurately, computed from off-the-shelf 2.5D dept…

Cited by 0SourceScholar
2026

Mitigating Multimodal Hallucinations via Gradient-based Self-Reflection

CVPR 2026

Multimodal large language models (MLLMs) achieve strong performance across diverse tasks but remain prone to hallucinations, where outputs are not grounded in visual inputs. This issue can be attributed to two main biases: text-visual bias, the overreliance on prompts and prior outputs, and co-occur

Cited by 0SourceScholar
2026

Temporal-Consistent Video Restoration with Pre-trained Diffusion Models

AAAI 2026technical

Video restoration (VR) aims to recover high-quality videos from degraded ones. Although recent zero-shot VR methods using pre-trained diffusion models (DMs) show good promise, they suffer from approximation errors during reverse diffusion and insufficient temporal consistency. Moreover, dealing with

Cited by 0SourcePDFScholar
2026

Towards Realistic and Consistent Orbital Video Generation via 3D Foundation Priors

CVPR 2026

We present a novel method for generating geometrically realistic and consistent orbital videos from a single image of an object. Existing video generation works mostly rely on pixel-wise attention to enforce view consistency across frames. However, such mechanism does not impose sufficient constrain

Cited by 0SourceScholar
2025

FRESA: Feedforward Reconstruction of Personalized Skinned Avatars from Few Images

CVPR 2025highlight

We present a novel method for reconstructing personalized 3D human avatars with realistic animation from only a few images. Due to the large variations in body shapes, poses, and cloth types, existing methods mostly require hours of per-subject optimization during inference, which limits their pract…

2025

GS-2DGS: Geometrically Supervised 2DGS for Reflective Object Reconstruction

CVPR 2025poster

3D modeling of highly reflective objects remains challenging due to strong view-dependent appearances. While previous SDF-based methods can recover high-quality meshes, they are often time-consuming and tend to produce over-smoothed surfaces. In contrast, 3D Gaussian Splatting (3DGS) offers the adva…

Cited by 0SourcePDFScholar
2025

Improving Cancer Gene Prediction by Enhancing Common Information Between the PPI Network and Gene Functional Association

AAAI 2025technical

Identifying cancer genes is crucial for treatment and understanding pathogenesis. Recent methods typically leverage protein-protein interaction (PPI) networks or gene functional association data from annotated gene sets. There may be some shared neighborhood structure information between these two t…

2025

Probability Density Geodesics in Image Diffusion Latent Space

CVPR 2025poster

Diffusion models indirectly estimate the probability density over a data space, which can be used to study its structure. In this work, we show that geodesics can be computed in diffusion latent space, where the norm induced by the spatially-varying inner product is inversely proportional to the pro…

2025

SRSR: Enhancing Semantic Accuracy in Real-World Image Super-Resolution with Spatially Re-Focused Text-Conditioning

NeurIPS 2025poster

Existing diffusion-based super-resolution approaches often exhibit semantic ambiguities due to inaccuracies and incompleteness in their text conditioning, coupled with the inherent tendency for cross-attention to divert towards irrelevant pixels. These limitations can lead to semantic misalignment a…

Cited by 0SourceScholar
2024

"NeuSDFusion: A Spatial-Aware Generative Model for 3D Shape Completion, Reconstruction, and Generation"

ECCV 2024poster

"3D shape generation aims to produce innovative 3D content adhering to specific conditions and constraints. Existing methods often decompose 3D shapes into a sequence of localized components, treating each element in isolation without considering spatial consistency. As a result, these approaches ex…

2024

Adapting Fine-Grained Cross-View Localization to Areas without Fine Ground Truth

ECCV 2024poster

"Given a ground-level query image and a geo-referenced aerial image that covers the query’s local surroundings, fine-grained cross-view localization aims to estimate the location of the ground camera inside the aerial image. Recent works have focused on developing advanced networks trained with accu…

2024

Advancing Virtual Reality Interaction: A Ring-Shaped Controller and Pose Tracking

ICRA 2024poster

Ensuring robust tracking of controllers’ movement is critical for human-robot interaction in virtual reality (VR) scenarios. This paper proposes a robust tracking algorithm based on a novel wearable ring-shaped controller equipped with an inertial measurement unit (IMU) and a light-emitting diode (L…

Cited by 0SourceScholar
2024

Alice Benchmarks: Connecting Real World Re-Identification with the Synthetic

ICLR 2024poster

For object re-identification (re-ID), learning from synthetic data has become a promising strategy to cheaply acquire large-scale annotated datasets and effective models, with few privacy concerns. Many interesting research problems arise from this strategy, e.g., how to reduce the domain gap betwee…

Cited by 0SourcePDFScholar
2024

ConsistNet: Enforcing 3D Consistency for Multi-view Images Diffusion

CVPR 2024poster

Given a single image of a 3D object this paper proposes a novel method (named ConsistNet) that can generate multiple images of the same object as if they are captured from different viewpoints while the 3D (multi-view) consistencies among those multiple generated images are effectively exploited. Ce…

2024

Increasing SLAM Pose Accuracy by Ground-to-Satellite Image Registration

ICRA 2024poster

Vision-based localization for autonomous driving has been of great interest among researchers. When a pre-built 3D map is not available, the techniques of visual simultaneous localization and mapping (SLAM) are typically adopted. Due to error accumulation, visual SLAM (vSLAM) usually suffers from lo…

Cited by 6SourcecodeScholar
2024

LAM3D: Large Image-Point Clouds Alignment Model for 3D Reconstruction from Single Image

NeurIPS 2024poster

Large Reconstruction Models have made significant strides in the realm of automated 3D content generation from single or multiple input images. Despite their success, these models often produce 3D meshes with geometric inaccuracies, stemming from the inherent challenges of deducing 3D shapes solely…

Cited by 3SourcePDFScholar
2024

MAVIS: Multi-Camera Augmented Visual-Inertial SLAM using SE2(3) Based Exact IMU Pre-integration

ICRA 2024poster

We present a novel optimization-based Visual-Inertial SLAM system designed for multiple partially over-lapped camera systems, named MAVIS. Our framework fully exploits the benefits of wide field-of-view from multi-camera systems, and the metric scale measurements provided by an inertial measurement…

Cited by 14SourcecodeScholar
2024

Prompting Future Driven Diffusion Model for Hand Motion Prediction

ECCV 2024poster

"Hand motion prediction from both first- and third-person perspectives is vital for enhancing user experience in AR/VR and ensuring safe remote robotic arm control. Previous works typically focus on predicting hand motion trajectories or human body motion, with direct hand motion prediction remainin…

Cited by 7SourcePDFScholar
2024

R$^2$-Gaussian: Rectifying Radiative Gaussian Splatting for Tomographic Reconstruction

NeurIPS 2024poster

3D Gaussian splatting (3DGS) has shown promising results in image rendering and surface reconstruction. However, its potential in volumetric reconstruction tasks, such as X-ray computed tomography, remains under-explored. This paper introduces R$^2$-Gaussian, the first 3DGS-based framework for spars…

2024

RGB-based Category-level Object Pose Estimation via Decoupled Metric Scale Recovery

ICRA 2024poster

While showing promising results, recent RGB-D camera-based category-level object pose estimation methods have restricted applications due to the heavy reliance on depth sensors. RGB-only methods provide an alternative to this problem yet suffer from inherent scale ambiguity stemming from monocular o…

Cited by 10SourcecodeScholar
2024

SemReg: Semantics Constrained Point Cloud Registration

ECCV 2024poster

"Despite the recent success of Transformers in point cloud registration, the cross-attention mechanism, while enabling point-wise feature exchange between point clouds, suffers from redundant feature interactions among semantically unrelated regions. Additionally, recent methods rely only on 3D info…

2024

StraightPCF: Straight Point Cloud Filtering

CVPR 2024poster

Point cloud filtering is a fundamental 3D vision task which aims to remove noise while recovering the underlying clean surfaces. State-of-the-art methods remove noise by moving noisy points along stochastic trajectories to the clean surfaces. These methods often require regularization within the tra…

2024

Towards High-Quality 3D Motion Transfer with Realistic Apparel Animation

ECCV 2024poster

"Animating stylized characters to match a reference motion sequence is a highly demanded task in film and gaming industries. Existing methods mostly focus on rigid deformations of characters’ body, neglecting local deformations on the apparel driven by physical dynamics. They deform apparel the same…

2024

View From Above: Orthogonal-View aware Cross-view Localization

CVPR 2024poster

This paper presents a novel aerial-to-ground feature aggregation strategy tailored for the task of cross-view image-based geo-localization. Conventional vision-based methods heavily rely on matching ground-view image features with a pre-recorded image database often through establishing planar homog…

Cited by 5SourcePDFScholar
2024

Weakly-supervised Camera Localization by Ground-to-satellite Image Registration

ECCV 2024poster

"The ground-to-satellite image matching/retrieval was initially proposed for city-scale ground camera localization. This work addresses the problem of improving camera pose accuracy by ground-to-satellite image matching after a coarse location and orientation have been obtained, either from the city…

2023

A Rotation-Translation-Decoupled Solution for Robust and Efficient Visual-Inertial Initialization

CVPR 2023poster

We propose a novel visual-inertial odometry (VIO) initialization method, which decouples rotation and translation estimation, and achieves higher efficiency and better robustness. Existing loosely-coupled VIO-initialization methods suffer from poor stability of visual structure-from-motion (SfM), wh…

2023

Boosting 3-DoF Ground-to-Satellite Camera Localization Accuracy via Geometry-Guided Cross-View Transformer

ICCV 2023poster

Image retrieval-based cross-view localization methods often lead to very coarse camera pose estimation, due to the limited sampling density of the database satellite images. In this paper, we propose a method to increase the accuracy of a ground camera's location and orientation by estimating the re…

Cited by 35PDFcodeScholar
2023

CDA: A Contrastive Data Augmentation Method for Alzheimer’s Disease Detection

ACL 2023findings

Alzheimer’s Disease (AD) is a neurodegenerative disorder that significantly impacts a patient’s ability to communicate and organize language. Traditional methods for detecting AD, such as physical screening or neurological testing, can be challenging and time-consuming. Recent research has explored…

Cited by 9SourcePDFScholar
2023

CircNet: Meshing 3D Point Clouds with Circumcenter Detection

ICLR 2023poster

Reconstructing 3D point clouds into triangle meshes is a key problem in computational geometry and surface reconstruction. Point cloud triangulation solves this problem by providing edge information to the input points. Since no vertex interpolation is involved, it is beneficial to preserve sharp de…

2023

DeepSimHO: Stable Pose Estimation for Hand-Object Interaction via Physics Simulation

NeurIPS 2023poster

This paper addresses the task of 3D pose estimation for a hand interacting with an object from a single image observation. When modeling hand-object interaction, previous works mainly exploit proximity cues, while overlooking the dynamical nature that the hand must stably grasp the object to counter…

2023

Homography Guided Temporal Fusion for Road Line and Marking Segmentation

ICCV 2023poster

Reliable segmentation of road lines and markings is critical to autonomous driving. Our work is motivated by the observations that road lines and markings are (1) frequently occluded in the presence of moving vehicles, shadow, and glare and (2) highly structured with low intra-class shape variance a…

Cited by 5PDFcodeScholar
2023

Inverting the Imaging Process by Learning an Implicit Camera Model

CVPR 2023poster

Representing visual signals with implicit coordinate-based neural networks, as an effective replacement of the traditional discrete signal representation, has gained considerable popularity in computer vision and graphics. In contrast to existing implicit neural representations which focus on modell…

Cited by 14SourcePDFScholar
2023

MB-TaylorFormer: Multi-Branch Efficient Transformer Expanded by Taylor Formula for Image Dehazing

ICCV 2023poster

In recent years, Transformer networks are beginning to replace pure convolutional neural networks (CNNs) in the field of computer vision due to their global receptive field and adaptability to input. However, the quadratic computational complexity of softmax-attention limits the wide application in…

Cited by 124PDFcodeScholar
2023

MEGANE: Morphable Eyeglass and Avatar Network

CVPR 2023poster

Eyeglasses play an important role in the perception of identity. Authentic virtual representations of faces can benefit greatly from their inclusion. However, modeling the geometric and appearance interactions of glasses and the face of virtual representations of humans is challenging. Glasses and f…

Cited by 16SourcePDFScholar
2023

Privacy Assessment on Reconstructed Images: Are Existing Evaluation Metrics Faithful to Human Perception?

NeurIPS 2023spotlight

Hand-crafted image quality metrics, such as PSNR and SSIM, are commonly used to evaluate model privacy risk under reconstruction attacks. Under these metrics, reconstructed images that are determined to resemble the original one generally indicate more privacy leakage. Images determined as overall d…

Cited by 8SourcePDFScholar
2023

Seeing Through the Glass: Neural 3D Reconstruction of Object Inside a Transparent Container

CVPR 2023poster

In this paper, we define a new problem of recovering the 3D geometry of an object confined in a transparent enclosure. We also propose a novel method for solving this challenging problem. Transparent enclosures pose challenges of multiple light reflections and refractions at the interface between di…

2023

View Consistent Purification for Accurate Cross-View Localization

ICCV 2023poster

This paper proposes a fine-grained self-localization method for outdoor robotics that utilizes a flexible number of onboard cameras and readily accessible satellite images. The proposed method addresses limitations in existing cross-view localization methods that struggle to handle noise sources suc…

Cited by 5PDFScholar
2022

Align and Prompt: Video-and-Language Pre-Training With Entity Prompts

CVPR 2022poster

Video-and-language pre-training has shown promising improvements on various downstream tasks. Most previous methods capture cross-modal interactions with a transformer-based multimodal encoder, not fully addressing the misalignment between unimodal video and text features. Besides, learning fine-gra…

Cited by 235PDFcodeScholar
2022

Beyond Cross-View Image Retrieval: Highly Accurate Vehicle Localization Using Satellite Image

CVPR 2022poster

This paper addresses the problem of vehicle-mounted camera localization by matching a ground-level image with an overhead-view satellite map. Existing methods often treat this problem as cross-view image retrieval, and use learned deep features to match the ground-level query image to a partition (e…

Cited by 96PDFcodeScholar
2022

Blind Image Decomposition

ECCV 2022poster

"We propose and study a novel task named Blind Image Decomposition (BID), which requires separating a superimposed image into constituent underlying images in a blind setting, that is, both the source components involved in mixing as well as the mixing mechanism are unknown. For example, rain may co…

2022

Improving GAN Equilibrium by Raising Spatial Awareness

CVPR 2022poster

The success of Generative Adversarial Networks (GANs) is largely built upon the adversarial training between a generator (G) and a discriminator (D). They are expected to reach a certain equilibrium where D cannot distinguish the generated images from the real ones. However, such an equilibrium is r…

Cited by 39PDFScholar
2022

You Only Cut Once: Boosting Data Augmentation with a Single Cut

ICML 2022spotlight

We present You Only Cut Once (YOCO) for performing data augmentations. YOCO cuts one image into two pieces and performs data augmentations individually within each piece. Applying YOCO improves the diversity of the augmentation per sample and encourages neural networks to recognize objects from part…

2021

ARVo: Learning All-Range Volumetric Correspondence for Video Deblurring

CVPR 2021poster

Video deblurring models exploit consecutive frames to remove blurs from camera shakes and object motions. In order to utilize neighboring sharp patches, typical methods rely mainly on homography or optical flows to spatially align neighboring blurry frames. However, such explicit approaches are less…

Cited by 83PDFScholar
2021

Benchmarking Ultra-High-Definition Image Super-Resolution

ICCV 2021poster

Increasingly, modern mobile devices allow capturing images at Ultra-High-Definition (UHD) resolution, which includes 4K and 8K images. However, current single image super-resolution (SISR) methods focus on super-resolving images to ones with resolution up to high definition (HD) and ignore higher-re…

Cited by 38PDFScholar
2021

Deep Two-View Structure-From-Motion Revisited

CVPR 2021poster

Two-view structure-from-motion (SfM) is the cornerstone of 3D reconstruction and visual SLAM. Existing deep learning-based approaches formulate the problem in ways that are fundamentally ill-posed, relying on training data to overcome the inherent difficulties. In contrast, we propose a return to th…

Cited by 63PDFcodeScholar
2021

Dual Pixel Exploration: Simultaneous Depth Estimation and Image Restoration

CVPR 2021poster

The dual-pixel (DP) hardware works by splitting each pixel in half and creating an image pair in a single snapshot. Several works estimate depth/inverse depth by treating the DP pair as a stereo pair. However, dual-pixel disparity only occurs in image regions with the defocus blur. The heavy defocus…

Cited by 43PDFScholar
2021

Learning To Estimate Hidden Motions With Global Motion Aggregation

ICCV 2021poster

Occlusions pose a significant challenge to optical flow algorithms that rely on local evidences. We consider an occluded point to be one that is imaged in the first frame but not in the next, a slight overloading of the standard definition since it also includes points that move out-of-frame. Estima…

Cited by 408PDFcodeScholar
2021

Multi-View 3D Reconstruction of a Texture-Less Smooth Surface of Unknown Generic Reflectance

CVPR 2021poster

Recovering the 3D geometry of a purely texture-less object with generally unknown surface reflectance (e.g. nonLambertian) is regarded as a challenging task in multiview reconstruction. The major obstacle revolves around establishing cross-view correspondences where photometric constancy is violated…

Cited by 31PDFcodeScholar
2021

Rethinking Class Relations: Absolute-Relative Supervised and Unsupervised Few-Shot Learning

CVPR 2021poster

The majority of existing few-shot learning methods describe image relations with binary labels. However, such binary relations are insufficient to teach the network complicated real-world relations, due to the lack of decision smoothness. Furthermore, current few-shot learning models capture only th…

Cited by 81PDFScholar
2020

Beyond Monocular Deraining: Stereo Image Deraining via Semantic Understanding

ECCV 2020poster

Rain is a common natural phenomenon. Taking images in the rain however often results in degraded quality of images, thus compromises the performance of many computer vision systems. Most existing de-rain algorithms use only one single input image and aim to recover a clean image. Few work has exploi…

Cited by 54SourcePDFScholar
2020

Channel Attention Based Iterative Residual Learning for Depth Map Super-Resolution

CVPR 2020poster

Despite the remarkable progresses made in deep learning based depth map super-resolution (DSR), how to tackle real-world degradation in low-resolution (LR) depth maps remains a major challenge. Existing DSR model is generally trained and tested on synthetic dataset, which is very different from what…

Cited by 103PDFScholar
2020

Deep Novel View Synthesis from Colored 3D Point Clouds

ECCV 2020poster

We propose a new deep neural network which takes a colored 3D point cloud of a scene, and directly synthesizes a photo-realistic image from an arbitrary viewpoint. Key contributions of this work include a deep point feature extraction module, an image synthesis module, and an image refinement module…

2020

Displacement-Invariant Matching Cost Learning for Accurate Optical Flow Estimation

NeurIPS 2020poster

Learning matching costs has been shown to be critical to the success of the state-of-the-art deep stereo matching methods, in which 3D convolutions are applied on a 4D feature volume to learn a 3D cost volume. However, this mechanism has never been employed for the optical flow task. This is mainly…

2020

End-to-end Learning for Inter-Vehicle Distance and Relative Velocity Estimation in ADAS with a Monocular Camera

ICRA 2020poster

Inter-vehicle distance and relative velocity estimations are two basic functions for any ADAS (Advanced driver-assistance systems). In this paper, we propose a monocular camera based inter-vehicle distance and relative velocity estimation method based on end-to-end training of a deep neural network.…

Cited by 26SourceScholar
2020

Few-shot Action Recognition with Permutation-invariant Attention

ECCV 2020poster

Many few-shot learning models focus on recognising images. In contrast, we tackle a challenging task of few-shot action recognition from videos. We build on a C3D encoder for spatio-temporal video blocks to capture short-range action patterns. Such encoded blocks are aggregated by permutation-invari…

Cited by 219SourcePDFScholar
2020

Globally Optimal Relative Pose Estimation for Camera on a Selfie Stick

ICRA 2020poster

Taking selfies has become a photographic trend nowadays. We envision the emergence of the "video selfie" capturing a short continuous video clip (or burst photography) of the user, themselves. A selfie stick is usually used, whereby a camera is mounted on a stick for taking selfie photos. In this sc…

Cited by 2SourceScholar
2020

Hierarchical Neural Architecture Search for Deep Stereo Matching

NeurIPS 2020poster

To reduce the human efforts in neural network design, Neural Architecture Search (NAS) has been applied with remarkable success to various high-level vision tasks such as classification and semantic segmentation. The underlying idea for the NAS algorithm is straightforward, namely, to allow the netw…

2020

Joint 3D Instance Segmentation and Object Detection for Autonomous Driving

CVPR 2020poster

Currently, in Autonomous Driving (AD), most of the 3D object detection frameworks (either anchor- or anchor-free-based) consider the detection as a Bounding Box (BBox) regression problem. However, this compact representation is not sufficient to explore all the information of the objects. To tackle…

Cited by 132PDFScholar
2020

Reliable frame-to-frame motion estimation for vehicle-mounted surround-view camera systems

ICRA 2020poster

Modern vehicles are often equipped with a surround-view multi-camera system. The current interest in autonomous driving invites the investigation of how to use such systems for a reliable estimation of relative vehicle displacement. Existing camera pose algorithms either work for a single camera, ma…

Cited by 12SourceScholar
2020

TSPNet: Hierarchical Feature Learning via Temporal Semantic Pyramid for Sign Language Translation

NeurIPS 2020poster

Sign language translation (SLT) aims to interpret sign video sequences into text-based natural language sentences. Sign videos consist of continuous sequences of sign gestures with no clear boundaries in between. Existing SLT models usually represent sign visual features in a frame-wise manner so as…

2020

Transferring Cross-Domain Knowledge for Video Sign Language Recognition

CVPR 2020oral

Word-level sign language recognition (WSLR) is a fundamental task in sign language interpretation. It requires models to recognize isolated sign words from videos. However, annotating WSLR data needs expert knowledge, thus limiting WSLR dataset acquisition. On the contrary, there are abundant subtit…

Cited by 164PDFScholar
2020

Where Am I Looking At? Joint Location and Orientation Estimation by Cross-View Matching

CVPR 2020poster

Cross-view geo-localization is the problem of estimating the position and orientation (latitude, longitude and azimuth angle) of a camera at ground level given a large-scale database of geo-tagged aerial (eg., satellite) images. Existing approaches treat the task as a pure location estimation proble…

Cited by 213PDFcodeScholar
2019

ApolloCar3D: A Large 3D Car Instance Understanding Benchmark for Autonomous Driving

CVPR 2019poster

Autonomous driving has attracted remarkable attention from both industry and academia. An important task is to estimate 3D properties (e.g. translation, rotation and shape) of a moving or parked vehicle on the road. This task, while critical, is still under-researched in the computer vision communit…

Cited by 224PDFcodeScholar
2019

Spatial-Aware Feature Aggregation for Image based Cross-View Geo-Localization

NeurIPS 2019poster

In this paper, we develop a new deep network to explicitly address these inherent differences between ground and aerial views. We observe there exist some approximate domain correspondences between ground and aerial images. Specifically, pixels lying on the same azimuth direction in an aerial image…

2019

The Alignment of the Spheres: Globally-Optimal Spherical Mixture Alignment for Camera Pose Estimation

CVPR 2019poster

Determining the position and orientation of a calibrated camera from a single image with respect to a 3D model is an essential task for many applications. When 2D-3D correspondences can be obtained reliably, perspective-n-point solvers can be used to recover the camera pose. However, without the pos…

Cited by 42PDFScholar
2019

Unsupervised Deep Epipolar Flow for Stationary or Dynamic Scenes

CVPR 2019poster

Unsupervised deep learning for optical flow computation has achieved promising results. Most existing deep-net based methods rely on image brightness consistency and local smoothness constraint to train the networks. Their performance degrades at regions where repetitive textures or occlusions occ…

Cited by 83PDFScholar
2018

Fully Convolutional Neural Networks for Road Detection with Multiple Cues Integration

ICRA 2018poster

Road detection from images is a key task in autonomous driving. The recent advent of deep learning (and in particular, CNN or convolutional neural networks) has greatly improved the performance of road detection algorithms. In this paper, we show how to fuse multiple different cues under the same co…

Cited by 11SourceScholar
2018

Scalable Dense Non-Rigid Structure-From-Motion: A Grassmannian Perspective

CVPR 2018poster

This paper addresses the task of dense non-rigid structure-from-motion (NRSfM) using multiple images. State-of-the-art methods to this problem are often hurdled by scalability, expensive computations, and noisy measurements. Further, recent methods to NRSfM usually either assume a small number of sp…

Cited by 58SourcePDFScholar
2018

Semi-Dense 3D Reconstruction with a Stereo Event Camera

ECCV 2018poster

Event cameras are bio-inspired sensors that offer several advantages, such as low latency, high-speed and high dynamic range, to tackle challenging scenarios in computer vision. This paper presents a solution to the problem of 3D reconstruction from data captured by a stereo event-camera rig moving…

Cited by 196SourcePDFScholar
2017

"Maximizing Rigidity" Revisited: A Convex Programming Approach for Generic 3D Shape Reconstruction From Multiple Perspective Views

ICCV 2017poster

Rigid structure-from-motion (RSfM) and non-rigid structure-from-motion (NRSfM) have long been treated in the literature as separate (different) problems. Inspired by a previous work which solved directly for 3D scene structure by factoring the relative camera poses out, we revisit the principle of "…

Cited by 20PDFScholar
2017

Globally-Optimal Inlier Set Maximisation for Simultaneous Camera Pose and Feature Correspondence

ICCV 2017oral

Estimating the 6-DoF pose of a camera from a single image relative to a pre-computed 3D point-set is an important task for many computer vision applications. Perspective-n-Point (PnP) solvers are routinely used for camera pose estimation, provided that a good quality set of 2D-3D feature corresponde…

Cited by 73PDFScholar
2017

Monocular Dense 3D Reconstruction of a Complex Dynamic Scene From Two Perspective Frames

ICCV 2017poster

This paper proposes a new approach for monocular dense 3D reconstruction of a complex dynamic scene from two perspective frames. By applying superpixel oversegmentation to the image, we model a generically dynamic (hence non-rigid) scene with a piecewise planar and rigid approximation. In this way,…

Cited by 74PDFScholar
2017

Neural Aggregation Network for Video Face Recognition

CVPR 2017poster

This paper presents a Neural Aggregation Network (NAN) for video face recognition. The network takes a face video or face image set of a person with a variable number of face images as its input, and produces a compact, fixed-dimension feature representation for recognition. The whole network is com…

Cited by 495PDFScholar
2017

Semi-dense visual odometry for RGB-D cameras using approximate nearest neighbour fields

ICRA 2017poster

This paper presents a robust and efficient semidense visual odometry solution for RGB-D cameras. The core of our method is a 2D-3D ICP pipeline which estimates the pose of the sensor by registering the projection of a 3D semidense map of a reference frame with the 2D semi-dense region extracted in t…

Cited by 17SourceScholar
2016

Real-time rotation estimation for dense depth sensors in piece-wise planar environments

IROS 2016poster

Low-drift rotation estimation is a crucial part of any accurate odometry system. In this paper, we focus on the problem of 3D rotation estimation with dense depth sensors in environments that consist of piece-wise planar structures, such as corridors and office rooms. An efficient mean-shift paradig…

Cited by 12SourceScholar
2016

Robust Optical Flow Estimation of Double-Layer Images Under Transparency or Reflection

CVPR 2016poster

This paper deals with a challenging, frequently encountered, yet not properly investigated problem in two-frame optical flow estimation. That is, the input frames are compounds of two imaging layers -- one desired background layer of the scene, and one distracting, possibly moving layer due to trans…

Cited by 62PDFScholar
2015

Iteratively Reweighted Graph Cut for Multi-Label MRFs With Non-Convex Priors

CVPR 2015poster

While widely acknowledged as highly effective in computer vision, multi-label MRFs with non-convex priors are difficult to optimize. To tackle this, we introduce an algorithm that iteratively approximates the original energy with an appropriately weighted surrogate energy that is easier to minimize.…

Cited by 14SourcePDFScholar
2015

Shape Interaction Matrix Revisited and Robustified: Efficient Subspace Clustering With Corrupted and Incomplete Data

ICCV 2015poster

The Shape Interaction Matrix (SIM) is one of the earliest approaches to performing subspace clustering (i.e., separating points drawn from a union of subspaces). In this paper, we revisit the SIM and reveal its connections to several recent subspace clustering methods. Our analysis lets us derive a…

Cited by 90PDFcodeScholar