← Search

Shenglong Hu

7 accepted papers

2026

Generalizable Co-Salient Object Detection via Mixed Content-Style Modulation

CVPR 2026

This paper presents a generalizable CoSOD framework via mixed content-style modulation, termed CoMCS, to enhance the robustness of the model to unseen domains. The CoMCS, consisting of a mixed content modulator (MCM), a mixed style modulator (MSM), and a collaborative semantic contrast module (SCM),

Cited by 0SourceScholar
2025

Continuously Learning Video-level Object Tokens for Robust UAV tracking

ICASSP 2025accepted

Due to the dynamic changes in flight motion and viewpoint, the objects in unmanned aerial vehicle (UAV) tracking scenarios often suffer from drastic appearance variations. Existing UAV trackers often leverage a frame-level matching mechanism, which measures the appearance similarity between the obje…

Cited by 0SourceScholar
2025

Group-wise Semantic-enhanced Interaction Network for Remote Sensing Spatio-Temporal Fusion

ICASSP 2025accepted

Remote sensing spatio-temporal fusion (STF) aims at fusing temporally-dense coarse-resolution images and temporally-sparse fine-resolution images to reconstruct high spatio-temporal resolution images. Multi-band remote sensing images are often accepted as inputs for STF that have complementary chara…

Cited by 0SourceScholar
2025

Learning Deep Frequency Degradation Prior for Remote Sensing Spatio-temporal Fusion

ICASSP 2025accepted

Existing deep learning-based remote sensing spatiotemporal fusion (STF) relies on a data-driven paradigm without considering the degradation prior modeling from the coarseto fine-resolution images. This makes the learned model easy to overfit to the training dataset, resulting in poor domain general…

Cited by 0SourceScholar
2025

Learning Joint Appearance and Shape Co-Representations for Co-Saliency Detection

ICASSP 2025accepted

Existing leading Co-saliency Detection (CoD) framework aims to segment the co-salient objects by learning the consensus visual representation of the foreground objects. However, despite different categories, some distractors may have similar appearance to the co-salient objects, such as Apples vs. B…

Cited by 0SourceScholar
2025

Open-Vocabulary Saliency-Guided Progressive Refinement Network for Unsupervised Video Object Segmentation

ICASSP 2025accepted

Existing leading unsupervised video object segmentation (UVOS) paradigm often leverages a dual-stream architecture with motion and appearance branches, where only the motion cues from optical flow are used as a guide to locating the primary foreground objects. When suffering from challenging factors…

Cited by 0SourceScholar
2025

Spatio-Semantic Prompt guided Adaptive Segment Anything for Remote Sensing Change Detection

ICASSP 2025accepted

Existing leading remote sensing change detection (RSCD) often takes a semantic-agnostic learning paradigm, which uses a binary ground-truth mask as supervision for model training. Despite the demonstrated success, due to the intrinsic characteristic of extremely complicated scene changes in RS image…

Cited by 0SourceScholar