← Search

Gangwei Xu

15 accepted papers

2026

BAT: Learning Event-based Optical Flow with Bidirectional Adaptive Temporal Correlation

AAAI 2026technical

Event cameras deliver visual information characterized by a high dynamic range and high temporal resolution, offering significant advantages in estimating optical flow for complex lighting conditions and fast-moving objects. Current advanced optical flow methods for event cameras largely adopt estab

Cited by 0SourcePDFScholar
2026

Generalized Geometry Encoding Volume for Real-time Stereo Matching

AAAI 2026technical

Real-time stereo matching methods primarily focus on enhancing in-domain performance but often overlook the critical importance of generalization in real-world applications. In contrast, recent stereo foundation models leverage monocular foundation models (MFMs) to improve generalization, but typica

Cited by 0SourcePDFScholar
2026

PromptStereo: Zero-Shot Stereo Matching via Structure and Motion Prompts

CVPR 2026

Modern stereo matching methods have leveraged monocular depth foundation models to achieve superior zero-shot generalization performance. However, most existing methods primarily focus on extracting robust features for cost volume construction or disparity initialization. At the same time, the itera

Cited by 0SourcecodeScholar
2026

ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving

ICLR 2026poster

Recent studies have explored leveraging the world knowledge and cognitive capabilities of Vision-Language Models (VLMs) to address the long-tail problem in end-to-end autonomous driving. However, existing methods typically formulate trajectory planning as a language modeling task, where physical act…

Cited by 0SourcecodeScholar
2025

ACP-MVS: Efficient Multi-View Stereo with Attention-based Context Perception

IROS 2025

The core of Multi-View Stereo (MVS) is to find corresponding pixels in neighboring images. However, due to challenging regions in input images such as untextured areas, repetitive patterns, or reflective surfaces, existing methods struggle to find precise pixel correspondence therein, resulting in i

Cited by 0SourcecodeScholar
2025

BANet: Bilateral Aggregation Network for Mobile Stereo Matching

ICCV 2025poster

State-of-the-art stereo matching methods typically use costly 3D convolutions to aggregate a full cost volume, but their computational demands make mobile deployment challenging. Directly applying 2D convolutions for cost aggregation often results in edge blurring, detail loss, and mismatches in tex…

2025

FlowMamba: Learning Point Cloud Scene Flow with Global Motion Propagation

AAAI 2025technical

Scene flow methods based on deep learning have achieved impressive performance. However, current top-performing methods still struggle with ill-posed regions, such as extensive flat regions or occlusions, due to insufficient local evidence. In this paper, we propose a novel global-aware scene flow e…

Cited by 2SourcePDFScholar
2025

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency

NeurIPS 2025poster

We present Genesis, a unified world model for joint generation of multi-view driving videos and LiDAR sequences with spatio-temporal and cross-modal consistency. Genesis employs a two-stage architecture that integrates a DiT-based video diffusion model with 3D-VAE encoding, and a BEV-represented LiD…

Cited by 0SourceScholar
2025

Leveraging Consistent Spatio-Temporal Correspondence for Robust Visual Odometry

AAAI 2025technical

Recent approaches to VO have significantly improved performance by using deep networks to predict optical flow between video frames. However, existing methods still suffer from noisy and inconsistent flow matching, making it difficult to handle challenging scenarios and long-sequence estimation.To o…

2025

MonSter: Marry Monodepth to Stereo Unleashes Power

CVPR 2025highlight

Stereo matching recovers depth from image correspondences. Existing methods struggle to handle ill-posed regions with limited matching cues, such as occlusions and textureless areas. To address this, we propose MonSter, a novel method that leverages the complementary strengths of monocular depth est…

2025

Pixel-Perfect Depth with Semantics-Prompted Diffusion Transformers

NeurIPS 2025poster

This paper presents **Pixel-Perfect Depth**, a monocular depth estimation model based on pixel-space diffusion generation that produces high-quality, flying-pixel-free point clouds from estimated depth maps. Current generative depth estimation models fine-tune Stable Diffusion and achieve impressive…

Cited by 0SourcecodeScholar
2024

HDRFlow: Real-Time HDR Video Reconstruction with Large Motions

CVPR 2024poster

Reconstructing High Dynamic Range (HDR) video from image sequences captured with alternating exposures is challenging especially in the presence of large camera or object motion. Existing methods typically align low dynamic range sequences using optical flow or attention mechanism for deghosting. Ho…

Cited by 15SourcePDFScholar
2024

Selective-Stereo: Adaptive Frequency Information Selection for Stereo Matching

CVPR 2024highlight

Stereo matching methods based on iterative optimization like RAFT-Stereo and IGEV-Stereo have evolved into a cornerstone in the field of stereo matching. However these methods struggle to simultaneously capture high-frequency information in edges and low-frequency information in smooth regions due t…

2023

Iterative Geometry Encoding Volume for Stereo Matching

CVPR 2023poster

Recurrent All-Pairs Field Transforms (RAFT) has shown great potentials in matching tasks. However, all-pairs correlations lack non-local geometry knowledge and have difficulties tackling local ambiguities in ill-posed regions. In this paper, we propose Iterative Geometry Encoding Volume (IGEV-Stereo…

2022

Attention Concatenation Volume for Accurate and Efficient Stereo Matching

CVPR 2022poster

Stereo matching is a fundamental building block for many vision and robotics applications. An informative and concise cost volume representation is vital for stereo matching of high accuracy and efficiency. In this paper, we present a novel cost volume construction method which generates attention w…

Cited by 283PDFcodeScholar