← Search

Pan Ji

25 accepted papers

2024

"NeuSDFusion: A Spatial-Aware Generative Model for 3D Shape Completion, Reconstruction, and Generation"

ECCV 2024poster

"3D shape generation aims to produce innovative 3D content adhering to specific conditions and constraints. Existing methods often decompose 3D shapes into a sequence of localized components, treating each element in isolation without considering spatial consistency. As a result, these approaches ex…

2024

Advancing Virtual Reality Interaction: A Ring-Shaped Controller and Pose Tracking

ICRA 2024poster

Ensuring robust tracking of controllers’ movement is critical for human-robot interaction in virtual reality (VR) scenarios. This paper proposes a robust tracking algorithm based on a novel wearable ring-shaped controller equipped with an inertial measurement unit (IMU) and a light-emitting diode (L…

Cited by 0SourceScholar
2024

ConsistNet: Enforcing 3D Consistency for Multi-view Images Diffusion

CVPR 2024poster

Given a single image of a 3D object this paper proposes a novel method (named ConsistNet) that can generate multiple images of the same object as if they are captured from different viewpoints while the 3D (multi-view) consistencies among those multiple generated images are effectively exploited. Ce…

2024

LAM3D: Large Image-Point Clouds Alignment Model for 3D Reconstruction from Single Image

NeurIPS 2024poster

Large Reconstruction Models have made significant strides in the realm of automated 3D content generation from single or multiple input images. Despite their success, these models often produce 3D meshes with geometric inaccuracies, stemming from the inherent challenges of deducing 3D shapes solely…

Cited by 3SourcePDFScholar
2024

RGB-based Category-level Object Pose Estimation via Decoupled Metric Scale Recovery

ICRA 2024poster

While showing promising results, recent RGB-D camera-based category-level object pose estimation methods have restricted applications due to the heavy reliance on depth sensors. RGB-only methods provide an alternative to this problem yet suffer from inherent scale ambiguity stemming from monocular o…

Cited by 10SourcecodeScholar
2024

SciCode: A Research Coding Benchmark Curated by Scientists

NeurIPS 2024poster

Since language models (LMs) now outperform average humans on many challenging tasks, it is becoming increasingly difficult to develop challenging, high-quality, and realistic evaluations. We address this by examining LM capabilities to generate code for solving real scientific research problems. Inc…

Cited by 18SourcePDFScholar
2023

RIAV-MVS: Recurrent-Indexing an Asymmetric Volume for Multi-View Stereo

CVPR 2023poster

This paper presents a learning-based method for multi-view depth estimation from posed images. Our core idea is a "learning-to-optimize" paradigm that iteratively indexes a plane-sweeping cost volume and regresses the depth map via a convolutional Gated Recurrent Unit (GRU). Since the cost volume pl…

2022

Deformable VisTR: Spatio Temporal Deformable Attention for Video Instance Segmentation

ICASSP 2022accepted

Video instance segmentation (VIS) task requires classifying, segmenting, and tracking object instances over all frames in a video clip. Recently, VisTR [1] has been proposed as end-to-end transformer-based VIS framework, while demonstrating state-of-the-art performance. However, VisTR is slow to con…

Cited by 0SourceScholar
2022

GeoRefine: Self-Supervised Online Depth Refinement for Accurate Dense Mapping

ECCV 2022poster

"We present a robust and accurate depth refinement system, named GeoRefine, for geometrically-consistent dense mapping from monocular sequences. GeoRefine consists of three modules: a hybrid SLAM module using learning-based priors, an online depth refinement module leveraging self-supervision, and a…

Cited by 11SourcePDFScholar
2022

PlaneMVS: 3D Plane Reconstruction From Multi-View Stereo

CVPR 2022poster

We present a novel framework named PlaneMVS for 3D plane reconstruction from multiple input views with known camera poses. Most previous learning-based plane reconstruction methods reconstruct 3D planes from single images, which highly rely on single-view regression and suffer from depth scale ambig…

Cited by 49PDFcodeScholar
2021

Invertible Denoising Network: A Light Solution for Real Noise Removal

CVPR 2021poster

Invertible networks have various benefits for image denoising since they are lightweight, information-lossless, and memory-saving during back-propagation. However, applying invertible models to remove noise is challenging because the input is noisy, and the reversed output is clean, following two di…

Cited by 200PDFcodeScholar
2021

MonoIndoor: Towards Good Practice of Self-Supervised Monocular Depth Estimation for Indoor Environments

ICCV 2021poster

Self-supervised depth estimation for indoor environments is more challenging than its outdoor counterpart in at least the following two aspects: (i) the depth range of indoor sequences varies a lot across different frames, making it difficult for the depth network to induce consistent depth cues, wh…

Cited by 96PDFScholar
2020

Displacement-Invariant Matching Cost Learning for Accurate Optical Flow Estimation

NeurIPS 2020poster

Learning matching costs has been shown to be critical to the success of the state-of-the-art deep stereo matching methods, in which 3D convolutions are applied on a 4D feature volume to learn a 3D cost volume. However, this mechanism has never been employed for the optical flow task. This is mainly…

2020

Learning Monocular Visual Odometry via Self-Supervised Long-Term Modeling

ECCV 2020poster

Monocular visual odometry (VO) suffers severely from error accumulation during frame-to-frame pose estimation. In this paper, we present a self-supervised learning method for VO with special consideration for consistency over longer sequences. To this end, we model the long-term dependency in pose p…

2020

Pseudo RGB-D for Self-Improving Monocular SLAM and Depth Prediction

ECCV 2020poster

Classical monocular Simultaneous Localization And Mapping (SLAM) and the recently emerging convolutional neural networks (CNNs) for monocular depth prediction represent two largely disjoint approaches towards building a 3D map of the surrounding environment. In this paper, we demonstrate that the co…

2019

Learning Structure-And-Motion-Aware Rolling Shutter Correction

CVPR 2019oral

An exact method of correcting the rolling shutter (RS) effect requires recovering the underlying geometry, i.e. the scene structures and the camera motions between scanlines or between views. However, the multiple-view geometry for RS cameras is much more complicated than its global shutter (GS) cou…

Cited by 64PDFScholar
2019

Unsupervised Deep Epipolar Flow for Stationary or Dynamic Scenes

CVPR 2019poster

Unsupervised deep learning for optical flow computation has achieved promising results. Most existing deep-net based methods rely on image brightness consistency and local smoothness constraint to train the networks. Their performance degrades at regions where repetitive textures or occlusions occ…

Cited by 83PDFScholar
2017

"Maximizing Rigidity" Revisited: A Convex Programming Approach for Generic 3D Shape Reconstruction From Multiple Perspective Views

ICCV 2017poster

Rigid structure-from-motion (RSfM) and non-rigid structure-from-motion (NRSfM) have long been treated in the literature as separate (different) problems. Inspired by a previous work which solved directly for 3D scene structure by factoring the relative camera poses out, we revisit the principle of "…

Cited by 20PDFScholar
2015

Shape Interaction Matrix Revisited and Robustified: Efficient Subspace Clustering With Corrupted and Incomplete Data

ICCV 2015poster

The Shape Interaction Matrix (SIM) is one of the earliest approaches to performing subspace clustering (i.e., separating points drawn from a union of subspaces). In this paper, we revisit the SIM and reveal its connections to several recent subspace clustering methods. Our analysis lets us derive a…

Cited by 90PDFcodeScholar