← Search

Zimin Xia

11 accepted papers

2026

Fusing Satellite Imagery and Planimetric Maps for Cross-View Localization

ICRA 2026poster

Current cross-view localization methods predominantly rely on satellite imagery as the aerial modality. Although recent work explores planimetric maps (e.g., OpenStreetMap tiles), these approaches often lag in performance. Yet both modalities are widely available and possess complementary properties…

2026

Loc$^{2}$: Interpretable Cross-View Localization via Depth-Lifted Local Feature Matching

ICLR 2026poster

We propose an accurate and interpretable fine-grained cross-view localization method that estimates the 3 Degrees of Freedom (DoF) pose of a ground-level image by matching its local features with a reference aerial image. Unlike prior approaches that rely on global descriptors or bird’s-eye-view (BE…

Cited by 0SourceScholar
2026

Real-time 3D Object Detection with Inference-Aligned Learning

AAAI 2026technical

Real-time 3D object detection from point clouds is essential for dynamic scene understanding in applications such as augmented reality, robotics, and navigation. We introduce a novel Spatial-prioritized and Rank-aware 3D object detection (SR3D) framework for indoor point clouds, to bridge the gap be

Cited by 0SourcePDFScholar
2026

Regulating Rather than Constraining: Adaptive Guidance for Complex Spectral Reconstruction in Pansharpening

CVPR 2026

In remote sensing pansharpening, spectrally mixed regions, where the spectral interactions among adjacent land covers lead to highly inconsistent reconstruction patterns, remain the most challenging areas. Due to the complex spatial distribution and heterogeneous spectral characteristics of ground o

Cited by 0SourcecodeScholar
2026

Sat3DGen: Comprehensive Street-Level 3D Scene Generation from Single Satellite Image

ICLR 2026poster

Generating a street-level 3D scene from a single satellite image is a crucial yet challenging task. Current methods present a stark trade-off: geometry-colorization models achieve high geometric fidelity but are typically building-focused and lack semantic diversity. In contrast, proxy-based models…

Cited by 0SourcecodeScholar
2025

CoMatcher: Multi-View Collaborative Feature Matching

CVPR 2025poster

This paper proposes a multi-view collaborative matching strategy for reliable track construction in complex scenarios. We observe that the pairwise matching paradigms applied to image set matching often result in ambiguous estimation when the selected independent pairs exhibit significant occlusions…

Cited by 0SourcePDFScholar
2025

GeoDistill: Geometry-Guided Self-Distillation for Weakly Supervised Cross-View Localization

ICCV 2025poster

Cross-view localization, the task of estimating a camera's 3-degrees-of-freedom (3-DoF) pose by aligning ground-level images with aerial images, is crucial for large-scale outdoor applications like autonomous navigation and augmented reality. Existing methods often rely on fully supervised learning,…

2023

SliceMatch: Geometry-Guided Aggregation for Cross-View Pose Estimation

CVPR 2023poster

This work addresses cross-view camera pose estimation, i.e., determining the 3-Degrees-of-Freedom camera pose of a given ground-level image w.r.t. an aerial image of the local area. We propose SliceMatch, which consists of ground and aerial feature extractors, feature aggregators, and a pose predict…

2022

Visual Cross-View Metric Localization with Dense Uncertainty Estimates

ECCV 2022poster

"This work addresses visual cross-view metric localization for outdoor robotics. Given a ground-level color image and a satellite patch that contains the local surroundings, the task is to identify the location of the ground camera within the satellite patch. Related work addressed this task for ran…

2021

Cross-View Matching for Vehicle Localization by Learning Geographically Local Representations

RA-L 2021

Cross-view matching aims to learn a shared image representation between ground-level images and satellite or aerial images at the same locations. In robotic vehicles, matching a camera image to a database of geo-referenced aerial imagery can serve as a method for self-localization. However, existing

Cited by 27SourceScholar