← Search

Jin Xie

66 accepted papers

2026

3D-Fixer: Coarse-to-Fine In-place Completion for 3D Scenes from a Single Image

CVPR 2026

Compositional 3D scene generation from a single view requires the simultaneous recovery of scene layout and 3D assets. Existing approaches mainly fall into two categories: feed-forward generation methods and per-instance generation methods. The former directly predict 3D assets with explicit 6DoF po

Cited by 0SourcecodeScholar
2026

FUSER: Feed-Forward Multiview 3D Registration Transformer and SE(3)$^N$ Diffusion Refinement

CVPR 2026

Registration of multiview point clouds typically depends on extensive pairwise matching to build a pose graph for global synchronization, which is computationally expensive and ill-posed without holistic geometric constraints. In this paper, we propose FUSER, the first feed-forward multi-view regist

Cited by 0SourcecodeScholar
2026

Few-Shot Incremental 3D Object Detection in Dynamic Indoor Environments

CVPR 2026

Incremental 3D object perception is a critical step toward embodied intelligence in dynamic indoor environments. However, existing incremental 3D detection methods rely on extensive annotations of novel classes for satisfactory performance. To address this limitation, we propose FI3Det, a Few-shot I

Cited by 0SourcecodeScholar
2026

GEM: Generating LiDAR World Model via Deformable Mamba

CVPR 2026

World models, which simulate environmental dynamics and generate sensor observations, are gaining increasing attention in autonomous driving. However, progress in LiDAR-based world models has lagged behind those built on camera videos or occupancy data, primarily due to two core challenges: the inhe

Cited by 0SourcecodeScholar
2026

IntrinsicWeather: Controllable Weather Editing in Intrinsic Space

CVPR 2026

We present IntrinsicWeather, a diffusion-based framework for controllable weather editing in intrinsic space. Our framework includes two components based on diffusion priors: an inverse renderer that estimates material properties, scene geometry, and lighting as intrinsic maps from an input image, a

Cited by 0SourceScholar
2026

VGGT-Long: Chunk It, Loop It, Align It -- Pushing VGGT's Limits on Kilometer-Scale Long RGB Sequences

ICRA 2026poster

Foundation models for 3D vision have recently demonstrated remarkable capabilities in 3D perception. However, extending these models to large-scale RGB stream 3D reconstruction remains challenging due to memory limitations. In this work, we propose VGGT-Long, a simple yet effective system that pushe…

2026

WorldSplat: Gaussian-Centric Feed-Forward 4D Scene Generation for Autonomous Driving

ICLR 2026poster

Recent advances in driving-scene generation and reconstruction have demonstrated significant potential for enhancing autonomous driving systems by producing scalable and controllable training data. Existing generation methods primarily focus on synthesizing diverse and high-fidelity driving videos;…

Cited by 0SourceScholar
2025

AD-GS: Object-Aware B-Spline Gaussian Splatting for Self-Supervised Autonomous Driving

ICCV 2025poster

Modeling and rendering dynamic urban driving scenes is crucial for self-driving simulation. Current high-quality methods typically rely on costly manual object tracklet annotations, while self-supervised approaches fail to capture dynamic object motions accurately and decompose scenes properly, resu…

Cited by 0SourcePDFScholar
2025

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation

ICCV 2025poster

Contrastive Language-Image Pre-training (CLIP) exhibits strong zero-shot classification ability on image-level tasks, leading to the research to adapt CLIP for open-vocabulary semantic segmentation without training. The key is to improve spatial representation of image-level CLIP, such as replacing…

2025

GSRecon: Efficient Generalizable Gaussian Splatting for Surface Reconstruction from Sparse Views

ICCV 2025poster

Generalizable surface reconstruction aims to recover the surface the scene from a sparse set of images in a feed-forward manner. Existing volume rendering-based methods evaluate numerous points along camera rays to infer the geometry, resulting in inefficient reconstruction. Recently, 3D Gaussian Sp…

2025

NaviFormer: A Spatio-Temporal Context-Aware Transformer for Object Navigation

AAAI 2025technical

Learning discriminative state representations of agents, encompassing the spatial layout and temporal pose trajectory, is essential for effective navigation decisions. However, existing approaches often rely on simplistic plain networks for navigation information fusion, overlooking the complex long…

2025

SSLFusion: Scale and Space Aligned Latent Fusion Model for Multimodal 3D Object Detection

AAAI 2025technical

Multimodal 3D object detection based on deep neural networks has indeed made significant progress. However, it still faces challenges due to the misalignment of scale and spatial information between features extracted from 2D images and those derived from 3D point clouds. Existing methods usually ag…

2025

SVG-IR: Spatially-Varying Gaussian Splatting for Inverse Rendering

CVPR 2025poster

Reconstructing 3D assets from images, known as inverse rendering (IR), remains a challenging task due to its ill-posed nature. 3D Gaussian Splatting (3DGS) has demonstrated impressive capabilities for novel view synthesis (NVS) tasks. Methods apply it to relighting by separating radiance into BRDF p…

2025

VoxelSplat: Dynamic Gaussian Splatting as an Effective Loss for Occupancy and Flow Prediction

CVPR 2025poster

Recent advancements in camera-based occupancy prediction have focused on the simultaneous prediction of 3D semantics and scene flow, a task that presents significant challenges due to specific difficulties, e.g., occlusions and unbalanced dynamic environments. In this paper, we analyze these challen…

2025

WeatherGen: A Unified Diverse Weather Generator for LiDAR Point Clouds via Spider Mamba Diffusion

CVPR 2025poster

3D scene perception demands a large amount of adverse-weather LiDAR data, yet the cost of LiDAR data collection presents a significant scaling-up challenge. To this end, a series of LiDAR simulators have been proposed. Yet, they can only simulate a single adverse weather with a single physical model…

2025

Zero-shot RGB-D Point Cloud Registration with Pre-trained Large Vision Model

CVPR 2025poster

This paper introduces ZeroMatch, a novel zero-shot RGB-D point cloud registration framework, aimed at achieving robust 3D matching on unseen data without any task-specific training. Our core idea is to utilize the powerful zero-shot image representation of Stable Diffusion, achieved through extensiv…

Cited by 0SourcePDFScholar
2024

Grid4D: 4D Decomposed Hash Encoding for High-Fidelity Dynamic Gaussian Splatting

NeurIPS 2024poster

Recently, Gaussian splatting has received more and more attention in the field of static scene rendering. Due to the low computational overhead and inherent flexibility of explicit representations, plane-based explicit methods are popular ways to predict deformations for Gaussian-based dynamic scene…

Cited by 2SourcePDFScholar
2024

Masked Motion Prediction with Semantic Contrast for Point Cloud Sequence Learning

ECCV 2024poster

"Self-supervised representation learning on point cloud sequences is a challenging task due to the complex spatio-temporal structure. Most recent attempts aim to train the point cloud sequences representation model by reconstructing the point coordinates or designing frame-level contrastive learning…

2024

SED: A Simple Encoder-Decoder for Open-Vocabulary Semantic Segmentation

CVPR 2024poster

Open-vocabulary semantic segmentation strives to distinguish pixels into different semantic groups from an open set of categories. Most existing methods explore utilizing pre-trained vision-language models in which the key is to adopt the image-level model for pixel-level segmentation task. In this…

2024

SGNet: Salient Geometric Network for Point Cloud Registration

IROS 2024poster

Point Cloud Registration (PCR) is a critical and challenging task in computer vision and robotics. One of the primary difficulties in PCR is identifying salient and meaningful points that exhibit consistent semantic and geometric properties across different scans. Previous methods have encountered c…

Cited by 0SourceScholar
2024

SPGroup3D: Superpoint Grouping Network for Indoor 3D Object Detection

AAAI 2024technical

Current 3D object detection methods for indoor scenes mainly follow the voting-and-grouping strategy to generate proposals. However, most methods utilize instance-agnostic groupings, such as ball query, leading to inconsistent semantic information and inaccurate regression of the proposals. To this…

2023

C2BN: Cross-Modality and Cross-Scale Balance Network for Multi-Modal 3D Object Detection

ICASSP 2023accepted

Multi-modal 3D object detection that classifies and locates objects in 3D space by combining point-clouds captured by lidars and RGB images captured by cameras, serves as the basis for autonomous driving. Most of the existing methods aggregate features from point-clouds and images by plain element-w…

Cited by 0SourceScholar
2023

Center-Based Decoupled Point-cloud Registration for 6D Object Pose Estimation

ICCV 2023poster

In this paper, we propose a novel center-based decoupled point cloud registration framework for robust 6D object pose estimation in real-world scenarios. Our method decouples the translation from the entire transformation by predicting the object center and estimating the rotation in a center-aware…

Cited by 12PDFScholar
2023

Graph Matching Optimization Network for Point Cloud Registration

IROS 2023poster

Point Cloud Registration is a fundamental and challenging problem in 3D computer vision. Recent works often utilize geometric structure features in downsampled points (patches) to seek correspondences, then propagate these sparse patch correspondences to the dense level in the corresponding patches'…

Cited by 4SourceScholar
2023

Hard Patches Mining for Masked Image Modeling

CVPR 2023poster

Masked image modeling (MIM) has attracted much research attention due to its promising potential for learning scalable visual representations. In typical approaches, models usually focus on predicting specific contents of masked patches, and their performances are highly related to pre-defined mask…

2023

Multi-Mode Online Knowledge Distillation for Self-Supervised Visual Representation Learning

CVPR 2023poster

Self-supervised learning (SSL) has made remarkable progress in visual representation learning. Some studies combine SSL with knowledge distillation (SSL-KD) to boost the representation learning performance of small models. In this study, we propose a Multi-mode Online Knowledge Distillation method (…

Cited by 38SourcePDFScholar
2023

Robust Outlier Rejection for 3D Registration With Variational Bayes

CVPR 2023poster

Learning-based outlier (mismatched correspondence) rejection for robust 3D registration generally formulates the outlier removal as an inlier/outlier classification problem. The core for this to be successful is to learn the discriminative inlier/outlier feature representations. In this paper, we de…

2023

SE(3) Diffusion Model-based Point Cloud Registration for Robust 6D Object Pose Estimation

NeurIPS 2023poster

In this paper, we introduce an SE(3) diffusion model-based point cloud registration framework for 6D object pose estimation in real-world scenarios. Our approach formulates the 3D registration task as a denoising diffusion process, which progressively refines the pose of the source point cloud to ob…

Cited by 28SourcePDFScholar
2023

Semantics-Consistent Feature Search for Self-Supervised Visual Representation Learning

ICCV 2023poster

In contrastive self-supervised learning, the common way to learn discriminative representation is to pull different augmented "views" of the same image closer while pushing all other images further apart, which has been proven to be effective. However, it is unavoidable to construct undesirable view…

Cited by 7PDFcodeScholar
2022

3D Siamese Transformer Network for Single Object Tracking on Point Clouds

ECCV 2022poster

"Siamese network based trackers formulate 3D single object tracking as cross-correlation learning between point features of a template and a search area. Due to the large appearance variation between the template and search area during tracking, how to learn the robust cross correlation between them…

2022

Domain Disentangled Generative Adversarial Network for Zero-Shot Sketch-Based 3D Shape Retrieval

AAAI 2022technical

Sketch-based 3D shape retrieval is a challenging task due to the large domain discrepancy between sketches and 3D shapes. Since existing methods are trained and evaluated on the same categories, they cannot effectively recognize the categories that have not been used during training. In this paper,…

Cited by 27SourcePDFScholar
2022

Generative Subgraph Contrast for Self-Supervised Graph Representation Learning

ECCV 2022poster

"Contrastive learning has shown great promise in the field of graph representation learning. By manually constructing positive/negative samples, most graph contrastive learning methods rely on the vector inner product based similarity metric to distinguish the samples for graph representation. Howev…

2022

Globally Optimal Relative Pose Estimation for Multi-Camera Systems with Known Gravity Direction

ICRA 2022poster

Multiple-camera systems have been widely used in self-driving cars, robots, and smartphones. In addition, they are typically also equipped with IMUs (inertial measurement units). Using the gravity direction extracted from the IMU data, the y-axis of the body frame of the multi-camera system can be a…

Cited by 3SourceScholar
2022

Learning Superpoint Graph Cut for 3D Instance Segmentation

NeurIPS 2022accept

3D instance segmentation is a challenging task due to the complex local geometric structures of objects in point clouds. In this paper, we propose a learning-based superpoint graph cut method that explicitly learns the local geometric structures of the point cloud for 3D instance segmentation. Speci…

Cited by 17SourcePDFScholar
2022

PSTR: End-to-End One-Step Person Search With Transformers

CVPR 2022poster

We propose a novel one-step transformer-based person search framework, PSTR, that jointly performs person detection and re-identification (re-id) in a single architecture. PSTR comprises a person search-specialized (PSS) module that contains a detection encoder-decoder for person detection along wit…

Cited by 75PDFcodeScholar
2022

RA-Depth: Resolution Adaptive Self-Supervised Monocular Depth Estimation

ECCV 2022poster

"Existing self-supervised monocular depth estimation methods can get rid of expensive annotations and achieve promising results. However, these methods suffer from severe performance degradation when directly adopting a model trained on a fixed resolution to evaluate at other different resolutions.…

2022

Reliable Inlier Evaluation for Unsupervised Point Cloud Registration

AAAI 2022technical

Unsupervised point cloud registration algorithm usually suffers from the unsatisfied registration precision in the partially overlapping problem due to the lack of effective inlier evaluation. In this paper, we propose a neighborhood consensus based reliable inlier evaluation method for robust unsup…

2022

Unsupervised Domain Adaptation for Point Cloud Semantic Segmentation via Graph Matching

IROS 2022poster

Unsupervised domain adaptation for point cloud semantic segmentation has attracted great attention due to its effectiveness in learning with unlabeled data. Most of existing methods use global-level feature alignment to transfer the knowledge from the source domain to the target domain, which may ca…

Cited by 14SourcecodeScholar
2021

3D Siamese Voxel-to-BEV Tracker for Sparse Point Clouds

NeurIPS 2021poster

3D object tracking in point clouds is still a challenging problem due to the sparsity of LiDAR points in dynamic environments. In this work, we propose a Siamese voxel-to-BEV tracker, which can significantly improve the tracking performance in sparse 3D point clouds. Specifically, it consists of a S…

2021

Action Candidate Based Clipped Double Q-learning for Discrete and Continuous Action Tasks

AAAI 2021technical

Double Q-learning is a popular reinforcement learning algorithm in Markov decision process (MDP) problems. Clipped Double Q-learning, as an effective variant of Double Q-learning, employs the clipped double estimator to approximate the maximum expected action value. Due to the underestimation bias o…

2021

Planning with Learned Dynamic Model for Unsupervised Point Cloud Registration

IJCAI 2021poster

Point cloud registration is a fundamental problem in 3D computer vision. In this paper, we cast point cloud registration into a planning problem in reinforcement learning, which can seek the transformation between the source and target point clouds through trial and error. By modeling the point clou…

Cited by 12SourcePDFScholar
2021

Pyramid Point Cloud Transformer for Large-Scale Place Recognition

ICCV 2021poster

Recently, deep learning based point cloud descriptors have achieved impressive results in the place recognition task. Nonetheless, due to the sparsity of point clouds, how to extract discriminative local features of point clouds to efficiently form a global descriptor is still a challenging problem.…

Cited by 142PDFcodeScholar
2021

SSPC-Net: Semi-supervised Semantic 3D Point Cloud Segmentation Network

AAAI 2021technical

Point cloud semantic segmentation is a crucial task in 3D scene understanding. Existing methods mainly focus on employing a large number of annotated labels for supervised semantic segmentation. Nonetheless, manually labeling such large point clouds for the supervised segmentation task is time-consu…

2021

Sampling Network Guided Cross-Entropy Method for Unsupervised Point Cloud Registration

ICCV 2021poster

In this paper, by modeling the point cloud registration task as a Markov decision process, we propose an end-to-end deep model embedded with the cross-entropy method (CEM) for unsupervised 3D registration. Our model consists of a sampling network module and a differentiable CEM module. In our sampli…

Cited by 45PDFcodeScholar
2021

Superpoint Network for Point Cloud Oversegmentation

ICCV 2021poster

Superpoints are formed by grouping similar points with local geometric structures, which can effectively reduce the number of primitives of point clouds for subsequent point cloud processing. Existing superpoint methods mainly focus on employing clustering or graph partition to generate superpoints…

Cited by 42PDFcodeScholar
2020

BidNet: Binocular Image Dehazing Without Explicit Disparity Estimation

CVPR 2020poster

Heavy haze results in severe image degradation and thus hampers the performance of visual perception, object detection, etc. On the assumption that dehazed binocular images are superior to the hazy ones for stereo vision tasks such as 3D object detection and according to the fact that image haze is…

Cited by 82PDFScholar
2020

Cascaded Non-local Neural Network for Point Cloud Semantic Segmentation

IROS 2020poster

In this paper, we propose a cascaded non-local neural network for point cloud segmentation. The proposed network aims to build the long-range dependencies of point clouds for the accurate segmentation. Specifically, we develop a novel cascaded non-local module, which consists of the neighborhood-lev…

Cited by 27SourceScholar
2020

Count- and Similarity-aware R-CNN for Pedestrian Detection

ECCV 2020poster

Recent pedestrian detection methods generally rely on additional supervision, such as visible bounding-box annotations, to handle heavy occlusions. We propose an approach that leverages pedestrian count and proposal similarity information within a two-stage pedestrian detection framework. Both pedes…

2020

Progressive Point Cloud Deconvolution Generation Network

ECCV 2020poster

In this paper, we propose an effective point cloud generation method, which can generate multi-resolution point clouds of the same shape from a latent vector. Specifically, we develop a novel progressive deconvolution network with the learning-based bilateral interpolation. The learning-based bilate…

2019

Deep Sketch-Shape Hashing With Segmented 3D Stochastic Viewing

CVPR 2019poster

Sketch-based 3D shape retrieval has been extensively studied in recent works, most of which focus on improving the retrieval accuracy, whilst neglecting the efficiency. In this paper, we propose a novel framework for efficient sketch-based 3D shape retrieval, i.e., Deep Sketch-Shape Hashing (DSSH),…

Cited by 48PDFScholar
2019

Mask-Guided Attention Network for Occluded Pedestrian Detection

ICCV 2019poster

Pedestrian detection relying on deep convolution neural networks has made significant progress. Though promising results have been achieved on standard pedestrians, the performance on heavily occluded pedestrians remains far from satisfactory. The main culprits are intra-class occlusions involving o…

Cited by 252PDFcodeScholar
2017

Learning Barycentric Representations of 3D Shapes for Sketch-Based 3D Shape Retrieval

CVPR 2017poster

Retrieving 3D shapes with sketches is a challenging problem since 2D sketches and 3D shapes are from two heterogeneous domains, which results in large discrepancy between them. In this paper, we propose to learn barycenters of 2D projections of 3D shapes for sketch-based 3D shape retrieval. Specific…

Cited by 91PDFScholar
2015

DeepShape: Deep Learned Shape Descriptor for 3D Shape Matching and Retrieval

CVPR 2015poster

Complex geometric structural variations of 3D models usually pose great challenges in 3D shape matching and retrieval. In this paper, we propose a high-level shape feature learning scheme to extract deformation-insensitive feature via a novel discriminative deep auto-encoder. First, we developed a m…

Cited by 182SourcePDFScholar