← Search

Yihang Luo

8 accepted papers

2026

4RC: 4D Reconstruction via Conditional Querying Anytime and Anywhere

ICML 2026poster

We present 4RC, a unified feed-forward framework for 4D reconstruction from monocular videos. Unlike existing methods that typically decouple motion from geometry or produce limited 4D attributes, such as sparse trajectories or two-view scene flow, 4RC learns a holistic 4D representation that jointl…

Cited by 0SourceScholar
2026

OmniVGGT: Omni-Modality Driven Visual Geometry Grounded Transformer

CVPR 2026

General 3D foundation models have started to lead the trend of unifying diverse vision tasks, yet most assume RGB-only inputs and ignore readily available geometric cues (e.g., camera intrinsics, poses, and depth maps). To address this issue, we introduce OmniVGGT, a novel framework that can effecti

Cited by 0SourcecodeScholar
2026

STream3R: Scalable Sequential 3D Reconstruction with Causal Transformer

ICLR 2026poster

We present STream3R, a novel approach to 3D reconstruction that reformulates pointmap prediction as a decoder-only Transformer problem. Existing state-of-the-art methods for multi-view reconstruction either depend on expensive global optimization or rely on simplistic memory mechanisms that scale po…

Cited by 0SourcecodeScholar
2026

Towards Persistence: Learning Topological Constraints for Event-based Small Object Detection

CVPR 2026

Small object detection (SOD) plays a vital role in applications such as anti-UAV tasks, yet conventional image-based methods struggle in high-speed scenarios due to the limited frame rate. Event cameras offer a promising alternative by capturing spatiotemporal event streams with microsecond-level te

Cited by 0SourceScholar
2026

VLANeXt: Recipes for Building Strong VLA Models

ICML 2026poster

Following the rise of large foundation models, Vision–Language–Action models (VLAs) emerged, leveraging strong visual and language understanding for general-purpose policy learning. Yet, the current VLA landscape remains fragmented and exploratory. Although many groups have proposed their own VLA mo…

Cited by 0SourceScholar
2025

3DEnhancer: Consistent Multi-View Diffusion for 3D Enhancement

CVPR 2025poster

Despite advances in neural rendering, due to the scarcity of high-quality 3D datasets and the inherent limitations of multi-view diffusion models, view synthesis and 3D model generation are restricted to low resolutions with suboptimal multi-view consistency. In this study, we present a novel 3D enh…

2024

Upscale-A-Video: Temporal-Consistent Diffusion Model for Real-World Video Super-Resolution

CVPR 2024highlight

Text-based diffusion models have exhibited remarkable success in generation and editing showing great promise for enhancing visual content with their generative prior. However applying these models to video super-resolution remains challenging due to the high demands for output fidelity and temporal…

Cited by 46SourcePDFScholar
2023

Nighttime Smartphone Reflective Flare Removal Using Optical Center Symmetry Prior

CVPR 2023highlight

Reflective flare is a phenomenon that occurs when light reflects inside lenses, causing bright spots or a "ghosting effect" in photos, which can impact their quality. Eliminating reflective flare is highly desirable but challenging. Many existing methods rely on manually designed features to detect…