← Search

LU FANG

28 accepted papers

2024

GigaHumanDet: Exploring Full-Body Detection on Gigapixel-Level Images

AAAI 2024technical

Performing person detection in super-high-resolution images has been a challenging task. For such a task, modern detectors, which usually encode a box using center and width/height, struggle with accuracy due to two factors: 1) Human characteristic: people come in various postures and the center wit…

Cited by 5SourcePDFScholar
2024

GigaTraj: Predicting Long-term Trajectories of Hundreds of Pedestrians in Gigapixel Complex Scenes

CVPR 2024poster

Pedestrian trajectory prediction is a well-established task with significant recent advancements. However existing datasets are unable to fulfill the demand for studying minute-level long-term trajectory prediction mainly due to the lack of high-resolution trajectory observation in the wide field of…

Cited by 3SourcePDFScholar
2024

OmniSeg3D: Omniversal 3D Segmentation via Hierarchical Contrastive Learning

CVPR 2024poster

Towards holistic understanding of 3D scenes a general 3D segmentation method is needed that can segment diverse objects without restrictions on object quantity or categories while also reflecting the inherent hierarchical structure. To achieve this we propose OmniSeg3D an omniversal segmentation met…

2024

SPECAT: SPatial-spEctral Cumulative-Attention Transformer for High-Resolution Hyperspectral Image Reconstruction

CVPR 2024poster

Compressive spectral image reconstruction is a critical method for acquiring images with high spatial and spectral resolution. Current advanced methods which involve designing deeper networks or adding more self-attention modules are limited by the scope of attention modules and the irrelevance of a…

2024

When Visual Grounding Meets Gigapixel-level Large-scale Scenes: Benchmark and Approach

CVPR 2024poster

Visual grounding refers to the process of associating natural language expressions with corresponding regions within an image. Existing benchmarks for visual grounding primarily operate within small-scale scenes with a few objects. Nevertheless recent advances in imaging technology have enabled the…

Cited by 5SourcePDFScholar
2024

XScale-NVS: Cross-Scale Novel View Synthesis with Hash Featurized Manifold

CVPR 2024poster

We propose XScale-NVS for high-fidelity cross-scale novel view synthesis of real-world large-scale scenes. Existing representations based on explicit surface suffer from discretization resolution or UV distortion while implicit volumetric representations lack scalability for large scenes due to the…

2023

Boosting Graph Contrastive Learning via Graph Contrastive Saliency

ICML 2023poster

Graph augmentation plays a crucial role in achieving good generalization for contrastive graph self-supervised learning. However, mainstream Graph Contrastive Learning (GCL) often favors random graph augmentations, by relying on random node dropout or edge perturbation on graphs. Random augmentation…

2023

Crowd3D: Towards Hundreds of People Reconstruction From a Single Image

CVPR 2023poster

Image-based multi-person reconstruction in wide-field large scenes is critical for crowd analysis and security alert. However, existing methods cannot deal with large scenes containing hundreds of people, which encounter the challenges of large number of people, large variations in human scale, and…

Cited by 13SourcePDFScholar
2023

DartBlur: Privacy Preservation With Detection Artifact Suppression

CVPR 2023poster

Nowadays, privacy issue has become a top priority when training AI algorithms. Machine learning algorithms are expected to benefit our daily life, while personal information must also be carefully protected from exposure. Facial information is particularly sensitive in this regard. Multiple datasets…

2023

PARF: Primitive-Aware Radiance Fusion for Indoor Scene Novel View Synthesis

ICCV 2023poster

This paper proposes a method for fast scene radiance field reconstruction with strong novel view synthesis performance and convenient scene editing functionality. The key idea is to fully utilize semantic parsing and primitive extraction for constraining and accelerating the radiance field reconstru…

Cited by 7PDFScholar
2023

RealGraph: A Multiview Dataset for 4D Real-world Context Graph Generation

ICCV 2023poster

In this paper, we propose a brand new scene understanding paradigm called "Context Graph Generation (CGG)", aiming at abstracting holistic semantic information in the complicated 4D world. The CGG task capitalizes on the calibrated multiview videos of a dynamic scene, and targets at recovering seman…

Cited by 1PDFcodeScholar
2023

SVQNet: Sparse Voxel-Adjacent Query Network for 4D Spatio-Temporal LiDAR Semantic Segmentation

ICCV 2023poster

LiDAR-based semantic perception tasks are critical yet challenging for autonomous driving. Due to the motion of objects and static/dynamic occlusion, temporal information plays an essential role in reinforcing perception by enhancing and completing single-frame knowledge. Previous approaches either…

Cited by 10PDFScholar
2022

ElasticMVS: Learning elastic part representation for self-supervised multi-view stereopsis

NeurIPS 2022accept

Self-supervised multi-view stereopsis (MVS) attracts increasing attention for learning dense surface predictions from only a set of images without onerous ground-truth 3D training data for supervision. However, existing methods highly rely on the local photometric consistency, which fails to identif…

Cited by 9SourcePDFScholar
2021

A2-FPN: Attention Aggregation Based Feature Pyramid Network for Instance Segmentation

CVPR 2021poster

Learning pyramidal feature representations is crucial for recognizing object instances at different scales. Feature Pyramid Network (FPN) is the classic architecture to build a feature pyramid with high-level semantics throughout. However, intrinsic defects in feature extraction and fusion inhibit F…

Cited by 122PDFScholar
2021

Data-Uncertainty Guided Multi-Phase Learning for Semi-Supervised Object Detection

CVPR 2021poster

In this paper, we delve into semi-supervised object detection where unlabeled images are leveraged to break through the upper bound of fully-supervised object detection models. Previous semi-supervised methods based on pseudo labels are severely degenerated by noise and prone to overfit to noisy lab…

Cited by 101PDFScholar
2021

LocalTrans: A Multiscale Local Transformer Network for Cross-Resolution Homography Estimation

ICCV 2021poster

Cross-resolution image alignment is a key problem in multiscale gigapixel photography, which requires to estimate homography matrix using images with large resolution gap. Existing deep homography methods concatenate the input images or features, neglecting the explicit formulation of correspondence…

Cited by 54PDFScholar
2020

EventCap: Monocular 3D Capture of High-Speed Human Motions Using an Event Camera

CVPR 2020oral

The high frame rate is a critical requirement for capturing fast human motions. In this setting, existing markerless image-based methods are constrained by the lighting requirement, the high data bandwidth and the consequent high computation overhead. In this paper, we propose EventCap -- the first…

Cited by 124PDFScholar
2020

PANDA: A Gigapixel-Level Human-Centric Video Dataset

CVPR 2020poster

We present PANDA, the first gigaPixel-level humAN-centric viDeo dAtaset, for large-scale, long-term, and multi-object visual analysis. The videos in PANDA were captured by a gigapixel camera and cover real-world scenes with both wide field-of-view ( 1 square kilometer area) and high-resolution detai…

Cited by 110PDFScholar
2020

RobustFusion: Human Volumetric Capture with Data-driven Visual Cues using a RGBD Camera

ECCV 2020poster

High-quality and complete 4D reconstruction of human activities is critical for immersive VR/AR experience, but it suffers from inherent self-scanning constraint and consequent fragile tracking under the monocular setting. In this paper, inspired by the huge potential of learning-based human modelin…

Cited by 106SourcePDFScholar
2019

PyTorch: An Imperative Style, High-Performance Deep Learning Library

NeurIPS 2019poster

Deep learning frameworks have often focused on either usability or speed, but not both. PyTorch is a machine learning library that shows that these two goals are in fact compatible: it was designed from first principles to support an imperative and Pythonic programming style that supports code as a…

2018

CrossNet: An End-to-end Reference-based Super Resolution Network using Cross-scale Warping

ECCV 2018poster

The Reference-based Super-resolution (RefSR) super-resolves a low-resolution (LR) image given an external high-resolution (HR) reference image, where the reference image and LR image share similar viewpoint but with significant resolution gap x8. Existing RefSR methods work in a cascaded way such as…

Cited by 270SourcePDFScholar
2018

HybridFusion: Real-Time Performance Capture Using a Single Depth Sensor and Sparse IMUs

ECCV 2018poster

We propose a light-weight and highly robust real-time human performance capture method based on a single depth camera and sparse inertial measurement units (IMUs). The proposed method combines non-rigid surface tracking and volumetric surface fusion to simultaneously reconstruct challenging motions,…

Cited by 112SourcePDFScholar
2017

SurfaceNet: An End-To-End 3D Neural Network for Multiview Stereopsis

ICCV 2017poster

This paper proposes an end-to-end learning framework for multiview stereopsis. We term the network SurfaceNet. It takes a set of images and their corresponding camera parameters as input and directly infers the 3D model. The key advantage of the framework is that both photo-consistency as well geome…

Cited by 496PDFcodeScholar