← Search

haozhe lin

11 accepted papers

2026

ElasticFormer: Detecting Objects in HRW Shots via Elastic Computing Vision Transformer

CVPR 2026

Recent advances in gigapixel-level imaging have brought High-Resolution Wide shots to the forefront of research. However, these images present significant challenges: extreme sparsity of foreground, gigapixel-level resolutions and diverse target counts. This makes traditional close-up detectors inac

Cited by 0SourceScholar
2026

GigaMoE: Sparsity-Guided Mixture of Experts for Efficient Gigapixel Object Detection

AAAI 2026technical

Object detection in High-Resolution Wide (HRW) shots, or gigapixel images, presents unique challenges due to extreme object sparsity and vast scale variations. State-of-the-art methods like SparseFormer have pioneered sparse processing by selectively focusing on important regions, yet they apply a u

Cited by 0SourcePDFScholar
2026

MetricHMSR: Metric Human Mesh and Scene Recovery from Monocular Images

CVPR 2026

We introduce MetricHMSR (Metric Human Mesh and Scene Recovery), a novel approach for metric human mesh and scene recovery from monocular images. Due to unrealistic assumptions in the camera model and inherent challenges in metric perception, existing approaches struggle to achieve human pose and met

Cited by 0SourcecodeScholar
2026

Ramba: Selective State-Space Models for Relational Deep Learning

ICML 2026poster

Relational Deep Learning aims to learn directly on multi-table databases, yet current methods face a fundamental tension: Transformers' quadratic complexity prohibits the large contexts relational data demands, while GNNs sacrifice global context for efficiency. We introduce Ramba, the first selecti…

Cited by 0SourceScholar
2026

ThinFormer: Channel Sparse Transformer for Efficient HRW Object Detection

IJCAI 2026

Object detection in high-resolution wide (HRW) shots presents unique challenges due to the extreme sparsity of objects and the variability in sparsity ratios across images. Conventional detectors, designed for close-up settings like MS COCO, struggle to generalize to these scenarios, leading to inef

Cited by 0Scholar
2024

GigaTraj: Predicting Long-term Trajectories of Hundreds of Pedestrians in Gigapixel Complex Scenes

CVPR 2024poster

Pedestrian trajectory prediction is a well-established task with significant recent advancements. However existing datasets are unable to fulfill the demand for studying minute-level long-term trajectory prediction mainly due to the lack of high-resolution trajectory observation in the wide field of…

Cited by 3SourcePDFScholar
2024

When Visual Grounding Meets Gigapixel-level Large-scale Scenes: Benchmark and Approach

CVPR 2024poster

Visual grounding refers to the process of associating natural language expressions with corresponding regions within an image. Existing benchmarks for visual grounding primarily operate within small-scale scenes with a few objects. Nevertheless recent advances in imaging technology have enabled the…

Cited by 5SourcePDFScholar
2023

Crowd3D: Towards Hundreds of People Reconstruction From a Single Image

CVPR 2023poster

Image-based multi-person reconstruction in wide-field large scenes is critical for crowd analysis and security alert. However, existing methods cannot deal with large scenes containing hundreds of people, which encounter the challenges of large number of people, large variations in human scale, and…

Cited by 13SourcePDFScholar
2023

DartBlur: Privacy Preservation With Detection Artifact Suppression

CVPR 2023poster

Nowadays, privacy issue has become a top priority when training AI algorithms. Machine learning algorithms are expected to benefit our daily life, while personal information must also be carefully protected from exposure. Facial information is particularly sensitive in this regard. Multiple datasets…

2023

RealGraph: A Multiview Dataset for 4D Real-world Context Graph Generation

ICCV 2023poster

In this paper, we propose a brand new scene understanding paradigm called "Context Graph Generation (CGG)", aiming at abstracting holistic semantic information in the complicated 4D world. The CGG task capitalizes on the calibrated multiview videos of a dynamic scene, and targets at recovering seman…

Cited by 1PDFcodeScholar