← Search

Shan Wang

10 accepted papers

2026

Joint Shadow Generation and Relighting via Light-Geometry Interaction Maps

ICLR 2026poster

We propose Light–Geometry Interaction (LGI) maps, a novel representation that encodes light-aware occlusion from monocular depth. Unlike ray tracing, which requires full 3D reconstruction, LGI captures essential light–shadow interactions reliably and accurately, computed from off-the-shelf 2.5D dept…

Cited by 0SourceScholar
2026

Mitigating Multimodal Hallucinations via Gradient-based Self-Reflection

CVPR 2026

Multimodal large language models (MLLMs) achieve strong performance across diverse tasks but remain prone to hallucinations, where outputs are not grounded in visual inputs. This issue can be attributed to two main biases: text-visual bias, the overreliance on prompts and prior outputs, and co-occur

Cited by 0SourceScholar
2024

Increasing SLAM Pose Accuracy by Ground-to-Satellite Image Registration

ICRA 2024poster

Vision-based localization for autonomous driving has been of great interest among researchers. When a pre-built 3D map is not available, the techniques of visual simultaneous localization and mapping (SLAM) are typically adopted. Due to error accumulation, visual SLAM (vSLAM) usually suffers from lo…

Cited by 6SourcecodeScholar
2024

Large Multilingual Models Pivot Zero-Shot Multimodal Learning across Languages

ICLR 2024spotlight

Recently there has been a significant surge in multimodal learning in terms of both image-to-text and text-to-image generation. However, the success is typically limited to English, leaving other languages largely behind. Building a competitive counterpart in other languages is highly challenging du…

2024

View From Above: Orthogonal-View aware Cross-view Localization

CVPR 2024poster

This paper presents a novel aerial-to-ground feature aggregation strategy tailored for the task of cross-view image-based geo-localization. Conventional vision-based methods heavily rely on matching ground-view image features with a pre-recorded image database often through establishing planar homog…

Cited by 5SourcePDFScholar
2023

Homography Guided Temporal Fusion for Road Line and Marking Segmentation

ICCV 2023poster

Reliable segmentation of road lines and markings is critical to autonomous driving. Our work is motivated by the observations that road lines and markings are (1) frequently occluded in the presence of moving vehicles, shadow, and glare and (2) highly structured with low intra-class shape variance a…

Cited by 5PDFcodeScholar
2023

Model Calibration in Dense Classification with Adaptive Label Perturbation

ICCV 2023poster

For safety-related applications, it is crucial to produce trustworthy deep neural networks whose prediction is associated with confidence that can represent the likelihood of correctness for subsequent decision-making. Existing dense binary classification models are prone to being over-confident. To…

Cited by 4PDFcodeScholar
2023

Satellite Image Based Cross-view Localization for Autonomous Vehicle

ICRA 2023poster

Existing spatial localization techniques for au-tonomous vehicles mostly use a pre-built 3D-HD map, often constructed using a survey-grade 3D mapping vehicle, which is not only expensive but also laborious. This paper shows that by using an off-the-shelf high-definition satellite image as a ready-to…

Cited by 25SourceScholar
2023

View Consistent Purification for Accurate Cross-View Localization

ICCV 2023poster

This paper proposes a fine-grained self-localization method for outdoor robotics that utilizes a flexible number of onboard cameras and readily accessible satellite images. The proposed method addresses limitations in existing cross-view localization methods that struggle to handle noise sources suc…

Cited by 5PDFScholar