← Search

Maojun Zhang

16 accepted papers

2026

AerialExtreMatch: A Benchmark for Extreme-View Image Matching and Localization

RA-L 2026

Image matching serves as a core component for UAV localization guided by satellite imagery. However, this task remains highly challenging due to the extreme viewpoint discrepancies between low-altitude UAV images and nadir-view satellite maps. Existing datasets primarily focus on ground-level or hig

Cited by 0SourcecodeScholar
2026

Diffusion-aided Extreme Video Compression with Lightweight Semantics Guidance

ICASSP 2026oral

Modern video codecs and learning-based approaches struggle for semantic reconstruction at extremely low bit-rates due to reliance on low-level spatiotemporal redundancies. Generative models, especially diffusion models, offer a new paradigm for video compression by leveraging high-level semantic und…

Cited by 0SourcePDFScholar
2026

LoD-Loc v3: Generalized Aerial Localization in Dense Cities using Instance Silhouette Alignment

CVPR 2026

We present LoD-Loc v3, a novel method for generalized aerial visual localization in dense urban environments. While prior work LoD-Loc v2 achieves localization through semantic building silhouette alignment with low-detail city models, it suffers from two key limitations: poor cross-scene generaliza

Cited by 0SourcecodeScholar
2026

Local Precise Refinement: A Dual-Gated Mixture-of-Experts for Enhancing Foundation Model Generalization against Spectral Shifts

CVPR 2026

Domain Generalization Semantic Segmentation (DGSS) in spectral remote sensing is severely challenged by spectral shifts across diverse acquisition conditions, which cause significant performance degradation for models deployed in unseen domains. While fine-tuning foundation models is a promising dir

Cited by 0SourceScholar
2026

NGC-GeoLoc: Neural GeoCoordinate Regression for GPS-Denied UAV Geo-Localization

RA-L 2026

Visual geo-localization without GPS prior remains a significant challenge for UAV navigation. Traditional retrievalbased methods suffer from scale and rotation variances between UAV images and satellite maps, and their inference speed degrades with increasing map size. To address these challenges, w

Cited by 0SourcecodeScholar
2026

PiLoT: Neural Pixel-to-3D Registration for UAV-based Ego and Target Geo-localization

CVPR 2026

We present PiLoT, a unified framework that tackles UAV-based ego and target geo-localization. Conventional approaches rely on decoupled pipelines that fuse GNSS and Visual-Inertial Odometry (VIO) for ego-pose estimation, and active sensors like laser rangefinders for target localization. However, th

Cited by 0SourcecodeScholar
2025

LoD-Loc v2: Aerial Visual Localization over Low Level-of-Detail City Models using Explicit Silhouette Alignment

ICCV 2025poster

We propose a novel method for aerial visual localization over low Level-of-Detail (LoD) city models. Previous wireframe-alignment-based method LoD-Loc [99] has shown promising localization results leveraging LoD models. However, LoD-Loc mainly relies on high-LoD (LoD3 or LoD2) city models, but the m…

2025

MIMO Channel as a Neural Function: Implicit Neural Representations for Extreme CSI Compression

ICASSP 2025accepted

Acquiring and utilizing accurate channel state information (CSI) is crucial for realizing the benefits of massive multiple-input multiple-output (MIMO) technology. Current CSI feedback approaches improve precision by employing advanced deep-learning methods to learn representative CSI features for a…

Cited by 0SourceScholar
2025

NTR-Gaussian: Nighttime Dynamic Thermal Reconstruction with 4D Gaussian Splatting Based on Thermodynamics

CVPR 2025poster

Thermal infrared imaging enables a non-invasive measurement of the surface temperature of objects with all-weather applicability. Leveraging such techniques for 3D reconstruction can accurately reflect the temperature distribution of a scene, thereby supporting applications such as building monitori…

Cited by 1SourcePDFScholar
2024

LoD-Loc: Aerial Visual Localization using LoD 3D Map with Neural Wireframe Alignment

NeurIPS 2024poster

We propose a new method named LoD-Loc for visual localization in the air. Unlike existing localization algorithms, LoD-Loc does not rely on complex 3D representations and can estimate the pose of an Unmanned Aerial Vehicle (UAV) using a Level-of-Detail (LoD) 3D map. LoD-Loc mainly achieves this goal…

2023

Deep Active Contours for Real-time 6-DoF Object Tracking

ICCV 2023poster

This paper solves the problem of real-time 6-DoF object tracking from an RGB video. Prior optimization-based methods optimize the object pose by aligning the projected model to the image based on handcrafted features, which are prone to suboptimal solutions. Recent learning-based methods use neural…

Cited by 15PDFcodeScholar
2023

Long-Term Visual Localization With Mobile Sensors

CVPR 2023poster

Despite the remarkable advances in image matching and pose estimation, image-based localization of a camera in a temporally-varying outdoor environment is still a challenging problem due to huge appearance disparity between query and reference images caused by illumination, seasonal and structural c…

2018

FishEyeRecNet: A Multi-Context Collaborative Deep Network for Fisheye Image Rectification

ECCV 2018poster

Images captured by sheye lenses violate the pinhole camera assumption and suer from distortions. Rectication of sheye images is therefore a crucial preprocessing step for many computer vision applications. In this paper, we propose an end-to-end multi-context collaborative deep network for removing…

Cited by 163SourcePDFScholar
2018

MoNet: Deep Motion Exploitation for Video Object Segmentation

CVPR 2018poster

In this paper, we propose a novel MoNet model to deeply exploit motion cues for boosting video object segmentation performance from two aspects, i.e., frame representation learning and segmentation refinement. Concretely, MoNet exploits computed motion cue (i.e., optical flow) to reinforce the repre…

Cited by 164SourcePDFScholar
2016

Fast anomaly detection in traffic surveillance video based on robust sparse optical flow

ICASSP 2016accepted

Fast abnormal events detection in video is important for intelligent analysis of video. This paper proposes a fast anomaly detection algorithm based on sparse optical flow. We improve the efficiency of optical flow computation with foreground mask and spacial sampling and increase the robustness of…

Cited by 0SourceScholar