← Search

Long Quan

36 accepted papers

2025

DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos

CVPR 2025highlight

Estimating video depth in open-world scenarios is challenging due to the diversity of videos in appearance, content motion, camera movement, and length. We present DepthCrafter, an innovative method for generating temporally consistent long depth sequences with intricate details for open-world video…

2025

Mani-GS: Gaussian Splatting Manipulation with Triangular Mesh

CVPR 2025poster

Neural 3D representations, such as Neural Radiation Fields (NeRF), excel at producing photorealistic rendering results but lack the flexibility for manipulation and editing which is crucial for content creation. However, manipulating NeRF is not highly controllable and requires a long training and i…

Cited by 10SourcePDFScholar
2025

Matrix3D: Large Photogrammetry Model All-in-One

CVPR 2025highlight

We present Matrix3D, a unified model that performs several photogrammetry subtasks, including pose estimation, depth prediction, and novel view synthesis using just the same model. Matrix3D utilizes a multi-modal diffusion transformer (DiT) to integrate transformations across several modalities, suc…

2024

ConTex-Human: Free-View Rendering of Human from a Single Image with Texture-Consistent Synthesis

CVPR 2024poster

In this work we propose a method to address the challenge of rendering a 3D human from a single image in a free-view manner. Some existing approaches could achieve this by using generalizable pixel-aligned implicit fields to reconstruct a textured mesh of a human or by employing a 2D diffusion model…

Cited by 6SourcePDFScholar
2024

Direct2.5: Diverse Text-to-3D Generation via Multi-view 2.5D Diffusion

CVPR 2024poster

Recent advances in generative AI have unveiled significant potential for the creation of 3D content. However current methods either apply a pre-trained 2D diffusion model with the time-consuming score distillation sampling (SDS) or a direct 3D diffusion model trained on limited 3D data losing genera…

Cited by 33SourcePDFScholar
2024

HiFi-123: Towards High-fidelity One Image to 3D Content Generation

ECCV 2024poster

"Recent advances in diffusion models have enabled 3D generation from a single image. However, current methods often produce suboptimal results for novel views, with blurred textures and deviations from the reference image, limiting their practical applications. In this paper, we introduce HiFi-123,…

Cited by 26SourcePDFScholar
2024

JointNet: Extending Text-to-Image Diffusion for Dense Distribution Modeling

ICLR 2024poster

We introduce JointNet, a novel neural network architecture for modeling the joint distribution of images and an additional dense modality (e.g., depth maps). JointNet is extended from a pre-trained text-to-image diffusion model, where a copy of the original network is created for the new dense moda…

Cited by 10SourcePDFScholar
2023

NeILF++: Inter-Reflectable Light Fields for Geometry and Material Estimation

ICCV 2023poster

We present a novel differentiable rendering framework for joint geometry, material, and lighting estimation from multi-view images. In contrast to previous methods which assume a simplified environment map or co-located flashlights, in this work, we formulate the lighting of a static scene as one ne…

Cited by 56PDFScholar
2022

ASpanFormer: Detector-Free Image Matching with Adaptive Span Transformer

ECCV 2022poster

"Generating robust and reliable correspondences across images is a fundamental task for a diversity of applications. To capture context at both global and local granularity, we propose ASpanFormer, a Transformer-based detector-free matcher that is built on hierarchical attention structure, adopting…

2022

Critical Regularizations for Neural Surface Reconstruction in the Wild

CVPR 2022poster

Neural implicit functions have recently shown promising results on surface reconstructions from multiple views. However, current methods still suffer from excessive time complexity and poor robustness when reconstructing unbounded or complex scenes. In this paper, we present RegSDF, which shows that…

Cited by 54PDFScholar
2022

NeILF: Neural Incident Light Field for Physically-Based Material Estimation

ECCV 2022poster

"We present a differentiable rendering framework for material and lighting estimation from multi-view images and a reconstructed geometry. In the framework, we represent scene lightings as the Neural Incident Light Field (NeILF) and material properties as the surface BRDF modelled by multi-layer per…

Cited by 111SourcePDFScholar
2021

Learning To Match Features With Seeded Graph Matching Network

ICCV 2021poster

Matching local features across images is a fundamental problem in computer vision. Targeting towards high accuracy and efficiency, we propose Seeded Graph Matching Network, a graph neural network with sparse structure to reduce redundant connectivity and learn compact representation. The network con…

Cited by 142PDFcodeScholar
2020

ASLFeat: Learning Local Features of Accurate Shape and Localization

CVPR 2020poster

This work focuses on mitigating two limitations in the joint learning of local feature detectors and descriptors. First, the ability to estimate the local shape (scale, orientation, etc.) of feature points is often neglected during dense feature extraction, while the shape-awareness is crucial to ac…

Cited by 379PDFcodeScholar
2020

BlendedMVS: A Large-Scale Dataset for Generalized Multi-View Stereo Networks

CVPR 2020poster

While deep learning has recently achieved great success on multi-view stereo (MVS), limited training data makes the trained model hard to be generalized to unseen scenarios. Compared with other computer vision tasks, it is rather difficult to collect a large-scale MVS dataset as it requires expensiv…

Cited by 534PDFcodeScholar
2020

D3Feat: Joint Learning of Dense Detection and Description of 3D Local Features

CVPR 2020oral

A successful point cloud registration often lies on robust establishment of sparse matches through discriminative 3D local features. Despite the fast evolution of learning-based 3D feature descriptors, little attention has been drawn to the learning of 3D feature detectors, even less for a joint lea…

Cited by 528PDFcodeScholar
2020

Joint Semantic Segmentation and Boundary Detection Using Iterative Pyramid Contexts

CVPR 2020poster

In this paper, we present a joint multi-task learning framework for semantic segmentation and boundary detection. The critical component in the framework is the iterative pyramid context module (PCM), which couples two tasks and stores the shared latent semantics to interact between the two tasks. F…

Cited by 169PDFScholar
2020

KFNet: Learning Temporal Camera Relocalization Using Kalman Filtering

CVPR 2020oral

Temporal camera relocalization estimates the pose with respect to each video frame in sequence, as opposed to one-shot relocalization which focuses on a still image. Even though the time dependency has been taken into account, current temporal relocalization methods still generally underperform the…

Cited by 99PDFcodeScholar
2020

Learning Discriminative Feature with CRF for Unsupervised Video Object Segmentation

ECCV 2020poster

In this paper, we introduce a novel network, called discriminative feature network (DFNet), to address the unsupervised video object segmentation task. To capture the inherent correlation among video frames, we learn K discriminative features (D-features) from the input image and reference images th…

Cited by 72SourcePDFScholar
2020

Self-Supervised Monocular 3D Face Reconstruction by Occlusion-Aware Multi-view Geometry Consistency

ECCV 2020poster

Recent learning-based approaches, in which models are trained by single-view images have shown promising results for monocular 3D face reconstruction, but they suffer from the ill-posed face pose and depth ambiguity issue. In contrast to previous works that only enforce 2D feature constraints, we pr…

2020

Stochastic Bundle Adjustment for Efficient and Scalable 3D Reconstruction

ECCV 2020poster

Current bundle adjustment solvers such as the Levenberg-Marquardt (LM) algorithm are limited by the bottleneck in solving the Reduced Camera System (RCS) whose dimension is proportional to the camera number. When the problem is scaled up, this step is neither efficient in computation nor manageable…

2019

Beyond Photometric Loss for Self-Supervised Ego-Motion Estimation

ICRA 2019poster

Accurate relative pose is one of the key components in visual odometry (VO) and simultaneous localization and mapping (SLAM). Recently, the self-supervised learning framework that jointly optimizes the relative pose and target image depth has attracted the attention of the community. Previous works…

Cited by 113SourcecodeScholar
2019

ContextDesc: Local Descriptor Augmentation With Cross-Modality Context

CVPR 2019oral

Most existing studies on learning local features focus on the patch-based descriptions of individual keypoints, whereas neglecting the spatial relations established from their keypoint locations. In this paper, we go beyond the local detail representation by introducing context awareness to augment…

Cited by 315PDFcodeScholar
2019

Cross-Atlas Convolution for Parameterization Invariant Learning on Textured Mesh Surface

CVPR 2019poster

We present a convolutional network architecture for direct feature learning on mesh surfaces through their atlases of texture maps. The texture map encodes the parameterization from 3D to 2D domain, rendering not only RGB values but also rasterized geometric features if necessary. Since the paramete…

Cited by 21PDFScholar
2019

Learning Two-View Correspondences and Geometry Using Order-Aware Network

ICCV 2019poster

Establishing correspondences between two images requires both local and global spatial context. Given putative correspondences of feature points in two views, in this paper, we propose Order-Aware Network, which infers the probabilities of correspondences being inliers and regresses the relative pos…

Cited by 468PDFcodeScholar
2019

Recurrent MVSNet for High-Resolution Multi-View Stereo Depth Inference

CVPR 2019poster

Deep learning has recently demonstrated its excellent performance for multi-view stereo (MVS). However, one major limitation of current learned MVS approaches is the scalability: the memory-consuming cost volume regularization makes the learned MVS hard to be applied to high-resolution scenes. In th…

Cited by 706PDFcodeScholar
2018

GeoDesc: Learning Local Descriptors by Integrating Geometry Constraints

ECCV 2018poster

Learned local descriptors based on Convolutional Neural Networks (CNNs) have achieved significant improvements on patch-based benchmarks, whereas not having demonstrated strong generalization ability on recent benchmarks of image-based 3D reconstruction. In this paper, we mitigate this limitation by…

Cited by 216SourcePDFScholar
2018

Learning and Matching Multi-View Descriptors for Registration of Point Clouds

ECCV 2018poster

Critical to the registration of point clouds is the establishment of a set of accurate correspondences between points in 3D space. The correspondence problem is generally addressed by the design of discriminative 3D local descriptors on the one hand, and the development of robust matching strategies…

Cited by 58SourcePDFScholar
2018

MVSNet: Depth Inference for Unstructured Multi-view Stereo

ECCV 2018poster

We present an end-to-end deep learning architecture for depth map inference from multi-view images. In the network, we first extract deep visual image features, and then build the 3D cost volume upon the reference camera frustum via the differentiable homography warping. Next, we apply 3D convolutio…

2018

Reconstructing Thin Structures of Manifold Surfaces by Integrating Spatial Curves

CVPR 2018poster

The manifold surface reconstruction in multi-view stereo often fails in retaining thin structures due to incomplete and noisy reconstructed point clouds. In this paper, we address this problem by leveraging spatial curves. The curve representation in nature is advantageous in modeling thin and elong…

Cited by 39SourcePDFScholar
2018

Very Large-Scale Global SfM by Distributed Motion Averaging

CVPR 2018poster

Global Structure-from-Motion (SfM) techniques have demonstrated superior efficiency and accuracy than the conventional incremental approach in many recent studies. This work proposes a divide-and-conquer framework to solve very large global SfM at the scale of millions of images. Specifically, we fi…

Cited by 183SourcePDFScholar
2017

Progressive Large Scale-Invariant Image Matching in Scale Space

ICCV 2017poster

The power of modern image matching approaches is still fundamentally limited by the abrupt scale changes in images. In this paper, we propose a scale-invariant image matching approach to tackling the very large scale variation of views. Drawing inspiration from the scale space theory, we start with…

Cited by 47PDFScholar
2015

Higher-Order CRF Structural Segmentation of 3D Reconstructed Surfaces

ICCV 2015poster

In this paper, we propose a structural segmentation algorithm to partition multi-view stereo reconstructed surfaces of large-scale urban environments into structural segments. Each segment corresponds to a structural component describable by a surface primitive of up to the second order. This segmen…

Cited by 18PDFScholar
2015

Joint Camera Clustering and Surface Segmentation for Large-Scale Multi-View Stereo

ICCV 2015poster

In this paper, we propose an optimal decomposition approach to large-scale multi-view stereo from an initial sparse reconstruction. The success of the approach depends on the introduction of surface-segmentation-based camera clustering rather than sparse-point-based camera clustering, which suffers…

Cited by 29PDFScholar