← Search

Kang Wu

7 accepted papers

2026

Label Enhancement via Cross-View Fusion and Mixed Graph Propagation

IJCAI 2026

Label Distribution Learning (LDL) effectively addresses label ambiguity by modeling the degree to which each label describes an instance. A key challenge in LDL is Label Enhancement (LE): recovering label distributions from logical labels. Existing LE methods typically treat logical labels as superv

Cited by 0Scholar
2026

SkySense-VITA: Towards Universal In-context Segmentation of Multi-modal Remote Sensing Imagery

CVPR 2026

While recent foundation models for remote sensing segmentation have shown notable progress, they still fall short in processing diverse multi-modal inputs, synergizing complementary prompt types, and leveraging semantic hierarchies. To address these limitations, we introduce SkySense-VITA, a unified

Cited by 0SourceScholar
2025

SkySense V2: A Unified Foundation Model for Multi-modal Remote Sensing

ICCV 2025poster

The multi-modal remote sensing foundation model (MM-RSFM) has significantly advanced various Earth observation tasks, such as urban planning, environmental monitoring, and natural disaster management. However, most existing approaches generally require the training of separate backbone networks for…

Cited by 0SourcePDFScholar
2025

SkySense-O: Towards Open-World Remote Sensing Interpretation with Vision-Centric Visual-Language Modeling

CVPR 2025poster

Open-world interpretation aims to accurately localize and recognize all objects within images by vision-language models (VLMs). While substantial progress has been made in this task for natural images, the advancements for remote sensing (RS) images still remain limited, primarily due to these two c…

2025

When Large Vision-Language Model Meets Large Remote Sensing Imagery: Coarse-to-Fine Text-Guided Token Pruning

ICCV 2025poster

Efficient vision-language understanding of large Remote Sensing Images (RSIs) is meaningful but challenging. Current Large Vision-Language Models (LVLMs) typically employ limited pre-defined grids to process images, leading to information loss when handling gigapixel RSIs. Conversely, using unlimite…

2024

Fine-Grained Scene Graph Generation via Sample-Level Bias Prediction

ECCV 2024poster

"Scene Graph Generation (SGG) aims to explore the relationships between objects in images and obtain scene summary graphs, thereby better serving downstream tasks. However, the long-tailed problem has adversely affected the scene graph’s quality. The predictions are dominated by coarse-grained relat…

2024

SkySense: A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation Imagery

CVPR 2024poster

Prior studies on Remote Sensing Foundation Model (RSFM) reveal immense potential towards a generic model for Earth Observation. Nevertheless these works primarily focus on a single modality without temporal and geo-context modeling hampering their capabilities for diverse tasks. In this study we pre…

Cited by 140SourcePDFScholar