← Search

Zhitong Xiong

9 accepted papers

2026

TerraScope: Pixel-Grounded Visual Reasoning for Earth Observation

CVPR 2026

Vision-language models (VLMs) have shown promise in earth observation (EO), yet they struggle with tasks that require grounding complex spatial reasoning in precise pixel-level visual representations. To address this problem, we introduce TerraScope, a unified VLM that delivers pixel-grounded geospa

Cited by 0SourceScholar
2025

REOBench: Benchmarking Robustness of Earth Observation Foundation Models

NeurIPS 2025poster

Earth observation foundation models have shown strong generalization across multiple Earth observation tasks, but their robustness under real-world perturbations remains underexplored. To bridge this gap, we introduce REOBench, the first comprehensive benchmark for evaluating the robustness of Earth…

Cited by 0SourcecodeScholar
2025

Towards a Unified Copernicus Foundation Model for Earth Vision

ICCV 2025poster

Advances in Earth observation (EO) foundation models have unlocked the potential of big satellite data to learn generic representations from space, benefiting a wide range of downstream applications crucial to our planet. However, most existing efforts remain limited to fixed spectral sensors, focus…

2024

Decoupling Common and Unique Representations for Multimodal Self-supervised Learning

ECCV 2024oral

"The increasing availability of multi-sensor data sparks wide interest in multimodal self-supervised learning. However, most existing approaches learn only common representations across modalities while ignoring intra-modal training and modality-unique representations. We propose Decoupling Common a…

2024

Representation Enhancement-Stabilization: Reducing Bias-Variance of Domain Generalization

ECCV 2024poster

"Domain Generalization (DG) focuses on enhancing the generalization of deep learning models trained on multiple source domains to adapt to unseen target domains. This paper explores DG through the lens of bias-variance decomposition, uncovering that test errors in DG predominantly arise from cross-d…

2022

Doubly Deformable Aggregation of Covariance Matrices for Few-Shot Segmentation

ECCV 2022poster

"Training semantic segmentation models with few annotated samples has great potential in various real-world applications. For the few-shot segmentation task, the main challenge is how to accurately measure the semantic correspondence between the support and query samples with limited training data.…

2020

KALM: Key Area Localization Mechanism for Abnormality Detection in Musculoskeletal Radiographs

ICASSP 2020accepted

Recently abnormality detection in musculoskeletal radio-graphs has attracted many attentions. For abnormality detection, it is crucial to locate the most important area in the musculoskeletal radiographs. To achieve this goal, we propose a key area localization mechanism (KALM) for abnormality detec…

Cited by 0SourceScholar