← Search

Kaihua Zhang

21 accepted papers

2026

Generalizable Co-Salient Object Detection via Mixed Content-Style Modulation

CVPR 2026

This paper presents a generalizable CoSOD framework via mixed content-style modulation, termed CoMCS, to enhance the robustness of the model to unseen domains. The CoMCS, consisting of a mixed content modulator (MCM), a mixed style modulator (MSM), and a collaborative semantic contrast module (SCM),

Cited by 0SourceScholar
2025

Continuously Learning Video-level Object Tokens for Robust UAV tracking

ICASSP 2025accepted

Due to the dynamic changes in flight motion and viewpoint, the objects in unmanned aerial vehicle (UAV) tracking scenarios often suffer from drastic appearance variations. Existing UAV trackers often leverage a frame-level matching mechanism, which measures the appearance similarity between the obje…

Cited by 0SourceScholar
2025

Easy-to-hard Instance-level Feature Fusion for Co-saliency Detection

ICASSP 2025accepted

Existing leading deep learning-based Co-saliency Detection (CoD) methods often learn the consensus features from the input image group without considering the complexity of each image. Despite the demonstrated success, the input images may contain hard samples with high complexity, e.g., those conta…

Cited by 0SourceScholar
2025

Group-wise Semantic-enhanced Interaction Network for Remote Sensing Spatio-Temporal Fusion

ICASSP 2025accepted

Remote sensing spatio-temporal fusion (STF) aims at fusing temporally-dense coarse-resolution images and temporally-sparse fine-resolution images to reconstruct high spatio-temporal resolution images. Multi-band remote sensing images are often accepted as inputs for STF that have complementary chara…

Cited by 0SourceScholar
2025

Learning Deep Frequency Degradation Prior for Remote Sensing Spatio-temporal Fusion

ICASSP 2025accepted

Existing deep learning-based remote sensing spatiotemporal fusion (STF) relies on a data-driven paradigm without considering the degradation prior modeling from the coarseto fine-resolution images. This makes the learned model easy to overfit to the training dataset, resulting in poor domain general…

Cited by 0SourceScholar
2025

Learning Joint Appearance and Shape Co-Representations for Co-Saliency Detection

ICASSP 2025accepted

Existing leading Co-saliency Detection (CoD) framework aims to segment the co-salient objects by learning the consensus visual representation of the foreground objects. However, despite different categories, some distractors may have similar appearance to the co-salient objects, such as Apples vs. B…

Cited by 0SourceScholar
2025

Open-Vocabulary Saliency-Guided Progressive Refinement Network for Unsupervised Video Object Segmentation

ICASSP 2025accepted

Existing leading unsupervised video object segmentation (UVOS) paradigm often leverages a dual-stream architecture with motion and appearance branches, where only the motion cues from optical flow are used as a guide to locating the primary foreground objects. When suffering from challenging factors…

Cited by 0SourceScholar
2025

Shifting Spotlight for Co-supervision: A Simple yet Efficient Single-branch Network to See Through Camouflage

ICASSP 2025accepted

Camouflaged object detection (COD) remains a challenging task in computer vision. Existing methods often resort to additional branches for edge supervision, incurring substantial computational costs. To address this, we propose the Co-Supervised Spotlight Shifting Network (CS<sup xmlns:mml="http://w…

Cited by 0SourceScholar
2025

Spatio-Semantic Prompt guided Adaptive Segment Anything for Remote Sensing Change Detection

ICASSP 2025accepted

Existing leading remote sensing change detection (RSCD) often takes a semantic-agnostic learning paradigm, which uses a binary ground-truth mask as supervision for model training. Despite the demonstrated success, due to the intrinsic characteristic of extremely complicated scene changes in RS image…

Cited by 0SourceScholar
2025

WeatherGen: A Unified Diverse Weather Generator for LiDAR Point Clouds via Spider Mamba Diffusion

CVPR 2025poster

3D scene perception demands a large amount of adverse-weather LiDAR data, yet the cost of LiDAR data collection presents a significant scaling-up challenge. To this end, a series of LiDAR simulators have been proposed. Yet, they can only simulate a single adverse weather with a single physical model…

2024

Generalizable Fourier Augmentation for Unsupervised Video Object Segmentation

AAAI 2024technical

The performance of existing unsupervised video object segmentation methods typically suffers from severe performance degradation on test videos when tested in out-of-distribution scenarios. The primary reason is that the test data in real- world may not follow the independent and identically distrib…

Cited by 6SourcePDFScholar
2024

Glance, Focus and Refinement Network for Remote Sensing Change Detection

ICASSP 2024accepted

Existing change detection (CD) methods often directly fuse the multi-level features from bi-temporal remote sensing images without discriminatively considering each pixel's importance. Despite the demonstrated success, unselectively mixing the features degrades the model's performance to effectively…

Cited by 0SourceScholar
2024

Segment Anything Model Guided Semantic Knowledge Learning For Remote Sensing Change Detection

ICASSP 2024accepted

Existing deep learning based remote sensing change detection (RSCD) methods only rely on binary ground-truth to guide the network learning while neglecting the useful semantic guidance. As a result, the network can be readily misled by irrelevant category changes, leading to degraded performance and…

Cited by 15SourceScholar
2024

Text2LiDAR: Text-guided LiDAR Point Clouds Generation via Equirectangular Transformer

ECCV 2024poster

"The complex traffic environment and various weather conditions make the collection of LiDAR data expensive and challenging. Achieving high-quality and controllable LiDAR data generation is urgently needed, controlling with text is a common practice, but there is little research in this field. To th…

2023

Co-Salient Object Detection With Uncertainty-Aware Group Exchange-Masking

CVPR 2023poster

The traditional definition of co-salient object detection (CoSOD) task is to segment the common salient objects in a group of relevant images. Existing CoSOD models by default adopt the group consensus assumption. This brings about model robustness defect under the condition of irrelevant images in…

Cited by 24SourcePDFScholar
2023

Distilling Cross-Temporal Contexts for Continuous Sign Language Recognition

CVPR 2023poster

Continuous sign language recognition (CSLR) aims to recognize glosses in a sign language video. State-of-the-art methods typically have two modules, a spatial perception module and a temporal aggregation module, which are jointly learned end-to-end. Existing results in [9,20,25,36] have indicated th…

Cited by 47SourcePDFScholar
2023

Group-Wise Co-Salient Object Detection with Siamese Transformers Via Brownian Distance Covariance Matching

ICASSP 2023accepted

Co-salient object detection (CoSOD) aims to discover and segment foreground targets in a group of images with the same semantic category. Existing mainstream approaches often employ convolutional neural networks (CNNs) to learn the semantic-invariant features from a group of images. Despite demonstr…

Cited by 0SourceScholar
2021

Deep Transport Network for Unsupervised Video Object Segmentation

ICCV 2021poster

The popular unsupervised video object segmentation methods fuse the RGB frame and optical flow via a two-stream network. However, they cannot handle the distracting noises in each input modality, which may vastly deteriorate the model performance. We propose to establish the correspondence between t…

Cited by 67PDFScholar
2021

DeepACG: Co-Saliency Detection via Semantic-Aware Contrast Gromov-Wasserstein Distance

CVPR 2021poster

The objective of co-saliency detection is to segment the co-occurring salient objects in a group of images. To address this task, we introduce a new deep network architecture via semantic-aware contrast Gromov-Wasserstein distance (DeepACG). We first adopt the Gromov-Wasserstein (GW) distance to bui…

Cited by 52PDFScholar
2020

Adaptive Graph Convolutional Network With Attention Graph Clustering for Co-Saliency Detection

CVPR 2020poster

Co-saliency detection aims to discover the common and salient foregrounds from a group of relevant images. For this task, we present a novel adaptive graph convolutional network with attention graph clustering (GCAGC). Three major contributions have been made, and are experimentally shown to have su…

Cited by 127PDFScholar
2019

Co-Saliency Detection via Mask-Guided Fully Convolutional Networks With Multi-Scale Label Smoothing

CVPR 2019poster

In image co-saliency detection problem, one critical issue is how to model the concurrent pattern of the co-salient parts, which appears both within each image and across all the relevant images. In this paper, we propose a hierarchical image co-saliency detection framework as a coarse to fine strat…

Cited by 110PDFScholar