← Search

Nan Xue

34 accepted papers

2026

FlowSSC: Universal Generative Monocular Semantic Scene Completion via One-Step Latent Diffusion

RA-L 2026

Semantic Scene Completion (SSC) from monocular RGB images is a fundamental yet challenging task due to the inherent ambiguity of inferring occluded 3D geometry from a single view. While feed-forward methods have made progress, they often struggle to generate plausible details in occluded regions and

Cited by 2SourceScholar
2026

Real-time 3D Object Detection with Inference-Aligned Learning

AAAI 2026technical

Real-time 3D object detection from point clouds is essential for dynamic scene understanding in applications such as augmented reality, robotics, and navigation. We introduce a novel Spatial-prioritized and Rank-aware 3D object detection (SR3D) framework for indoor point clouds, to bridge the gap be

Cited by 0SourcePDFScholar
2025

FLARE: Feed-forward Geometry, Appearance and Camera Estimation from Uncalibrated Sparse Views

CVPR 2025poster

We present FLARE, a feed-forward model designed to infer high-quality camera poses and 3D geometry from uncalibrated sparse-view images (i.e., as few as 2-8 inputs), which is a challenging yet practical setting in real-world applications. Our solution features a cascaded learning paradigm with camer…

Cited by 0SourcePDFScholar
2025

PLANA3R: Zero-shot Metric Planar 3D Reconstruction via Feed-forward Planar Splatting

NeurIPS 2025poster

This paper addresses metric 3D reconstruction of indoor scenes by exploiting their inherent geometric regularities with compact representations. Using planar 3D primitives -- a well-suited representation for man-made environments -- we introduce PLANA3R, a pose-free framework for metric $\underline{…

Cited by 0SourcecodeScholar
2025

Rectified Diffusion Guidance for Conditional Generation

CVPR 2025poster

Classifier-Free Guidance (CFG), which combines the conditional and unconditional score functions with two coefficients summing to one, serves as a practical technique for diffusion model sampling. Theoretically, however, denoising with CFG cannot be expressed as a reciprocal diffusion process, which…

2025

ScaleLSD: Scalable Deep Line Segment Detection Streamlined

CVPR 2025poster

This paper studies the problem of Line Segment Detection (LSD) for the characterization of line geometry in images, with the aim of learning a domain-agnostic robust LSD model that works well for any natural images. With the focus of scalable self-supervised learning of LSD, we revisit and streamlin…

2025

SpatialTrackerV2: Advancing 3D Point Tracking with Explicit Camera Motion

ICCV 2025poster

We present SpatialTrackerV2, a feed-forward 3D point tracking method for monocular videos. Going beyond modular pipelines built on off-the-shelf components for 3D tracking, our approach unifies the intrinsic connections between point tracking, monocular depth, and camera pose estimation into a high-…

Cited by 0SourcePDFScholar
2024

Multi-View Attentive Contextualization for Multi-View 3D Object Detection

CVPR 2024poster

We present Multi-View Attentive Contextualization (MvACon) a simple yet effective method for improving 2D-to-3D feature lifting in query-based multi-view 3D (MV3D) object detection. Despite remarkable progress witnessed in the field of query-based MV3D object detection prior art often suffers from e…

Cited by 2SourcePDFScholar
2024

NEAT: Distilling 3D Wireframes from Neural Attraction Fields

CVPR 2024poster

This paper studies the problem of structured 3D recon- struction using wireframes that consist of line segments and junctions focusing on the computation of structured boundary geometries of scenes. Instead of leveraging matching-based solutions from 2D wireframes (or line segments) for 3D wireframe…

2024

SpatialTracker: Tracking Any 2D Pixels in 3D Space

CVPR 2024highlight

Recovering dense and long-range pixel motion in videos is a challenging problem. Part of the difficulty arises from the 3D-to-2D projection process leading to occlusions and discontinuities in the 2D motion domain. While 2D motion can be intricate we posit that the underlying 3D motion can often be…

2024

Stratified Avatar Generation from Sparse Observations

CVPR 2024poster

Estimating 3D full-body avatars from AR/VR devices is essential for creating immersive experiences in AR/VR applications. This task is challenging due to the limited input from Head Mounted Devices which capture only sparse observations from the head and hands. Predicting the full-body avatars parti…

Cited by 4SourcePDFScholar
2023

ConDaFormer: Disassembled Transformer with Local Structure Enhancement for 3D Point Cloud Understanding

NeurIPS 2023poster

Transformers have been recently explored for 3D point cloud understanding with impressive progress achieved. A large number of points, over 0.1 million, make the global self-attention infeasible for point cloud data. Thus, most methods propose to apply the transformer in a local region, e.g., spheri…

2023

HGFormer: Hierarchical Grouping Transformer for Domain Generalized Semantic Segmentation

CVPR 2023poster

Current semantic segmentation models have achieved great success under the independent and identically distributed (i.i.d.) condition. However, in real-world applications, test data might come from a different domain than training data. Therefore, it is important to improve model robustness against…

2023

Level-S$^2$fM: Structure From Motion on Neural Level Set of Implicit Surfaces

CVPR 2023poster

This paper presents a neural incremental Structure-from-Motion (SfM) approach, Level-S2fM, which estimates the camera poses and scene geometry from a set of uncalibrated images by learning coordinate MLPs for the implicit surfaces and the radiance fields from the established keypoint correspondences…

2023

Monocular 3D Object Detection with Bounding Box Denoising in 3D by Perceiver

ICCV 2023poster

The main challenge of monocular 3D object detection is the accurate localization of 3D center. Motivated by a new and strong observation that this challenge can be remedied by a 3D-space local-grid search scheme in an ideal case, we propose a stage-wise approach, which combines the information flow…

Cited by 14PDFScholar
2023

Sat2Density: Faithful Density Learning from Satellite-Ground Image Pairs

ICCV 2023poster

This paper aims to develop an accurate 3D geometry representation of satellite images using satellite-ground image pairs. Our focus is on the challenging problem of 3D-aware ground-views synthesis from a satellite image. We draw inspiration from the density field representation used in volumetric ne…

Cited by 17PDFcodeScholar
2022

Learning Auxiliary Monocular Contexts Helps Monocular 3D Object Detection

AAAI 2022technical

Monocular 3D object detection aims to localize 3D bounding boxes in an input single 2D image. It is a highly challenging problem and remains open, especially when no extra information (e.g., depth, lidar and/or multi-frames) can be leveraged in training and/or inference. This paper proposes a simpl…

2022

Learning Local-Global Contextual Adaptation for Multi-Person Pose Estimation

CVPR 2022poster

This paper studies the problem of multi-person pose estimation in a bottom-up fashion. With a new and strong observation that the localization issue of the center-offset formulation can be remedied in a local-window search scheme in an ideal situation, we propose a multi-person pose estimation appro…

Cited by 48PDFcodeScholar
2022

Partial Wasserstein Adversarial Network for Non-rigid Point Set Registration

ICLR 2022poster

Given two point sets, the problem of registration is to recover a transformation that matches one set to the other. This task is challenging due to the presence of large number of outliers, the unknown non-rigid deformations and the large sizes of point sets. To obtain strong robustness against outl…

Cited by 5SourcePDFScholar
2022

Revisiting Document Image Dewarping by Grid Regularization

CVPR 2022poster

This paper addresses the problem of document image dewarping, which aims at eliminating the geometric distortion in document images for document digitization. Instead of designing a better neural network to approximate the optical flow fields between the inputs and outputs, we pursue the best readab…

Cited by 34PDFcodeScholar
2020

Holistically-Attracted Wireframe Parsing

CVPR 2020poster

This paper presents a fast and parsimonious parsing method to accurately and robustly detect a vectorized wireframe in an input image with a single forward pass. The proposed method is end-to-end trainable, consisting of three components: (i) line segment and junction proposal generation, (ii) line…

Cited by 136PDFcodeScholar
2019

Learning Attraction Field Representation for Robust Line Segment Detection

CVPR 2019poster

This paper presents a region-partition based attraction field dual representation for line segment maps, and thus poses the problem of line segment detection (LSD) as the region coloring problem. The latter is then addressed by learning deep convolutional neural networks (ConvNets) for accur…

Cited by 158PDFcodeScholar
2019

Learning RoI Transformer for Oriented Object Detection in Aerial Images

CVPR 2019poster

Object detection in aerial images is an active yet challenging task in computer vision because of the bird's-eye view perspective, the highly complex backgrounds, and the variant appearances of objects. Especially when detecting densely packed objects in aerial images, methods relying on horizontal…

Cited by 1326PDFcodeScholar