← Search

Shuang Wang

27 accepted papers

2026

FedMPT: Federated Multi-Label Prompt Tuning of Vision-Language Models

CVPR 2026

Multi-Label Recognition (MLR) based on Vision-Language Models (VLMs) aims to leverage their pre-trained knowledge to better adapt complex recognition scenarios, thereby enhancing model robustness. However, for realistic decentralized applications requiring federated learning, adapting VLMs to each c

Cited by 0SourceScholar
2026

Generalizable Knowledge Distillation from Vision Foundation Models for Semantic Segmentation

CVPR 2026

Knowledge distillation (KD) has been widely applied in semantic segmentation to compress large models, but conventional approaches primarily preserve in-domain accuracy while neglecting out-of-domain generalization, which is essential under distribution shifts. This limitation becomes more severe wi

Cited by 0SourcecodeScholar
2026

Learning What Matters Now: Dynamic Preference Inference under Contextual Shifts

ICLR 2026poster

Humans often juggle multiple, sometimes conflicting objectives and shift their priorities as circumstances change, rather than following a fixed objective function. In contrast, most computational decision-making and multi-objective RL methods assume static preference weights or a known scalar rewa…

Cited by 0SourcecodeScholar
2025

ChangeDiff: A Multi-Temporal Change Detection Data Generator with Flexible Text Prompts via Diffusion Model

AAAI 2025technical

Data-driven deep learning models have enabled tremendous progress in change detection (CD) with the support of pixel-level annotations. However, collecting diverse data and manually annotating them is costly, laborious, and knowledge-intensive. Existing generative methods for CD data synthesis show…

2025

City-Level Foreign Direct Investment Prediction with Tabular Learning on Judicial Data

IJCAI 2025

To advance the United Nations Sustainable Development Goal on promoting sustained, inclusive, and sustainable economic growth, foreign direct investment (FDI) plays a crucial role in catalyzing economic expansion and fostering innovation. Precise city-level FDI prediction is quite important for loca

Cited by 0SourcePDFScholar
2025

Domain-aware Category-level Geometry Learning Segmentation for 3D Point Clouds

ICCV 2025poster

Domain generalization in 3D segmentation is a critical challenge in deploying models to unseen environments. Current methods mitigate the domain shift by augmenting the data distribution of point clouds. However, the model learns global geometric patterns in point clouds while ignoring the category-…

2025

Feature Spectrum Learning for Remote Sensing Change Detection

CVPR 2025poster

Change detection (CD) holds significant implications for Earth observation, in which pseudo-changes between bitemporal images induced by imaging environmental factors are key challenges. Existing methods mainly regard pseudo-changes as a kind of style shift and alleviate it by transforming bitempora…

Cited by 0SourcePDFScholar
2025

FisherTune: Fisher-Guided Robust Tuning of Vision Foundation Models for Domain Generalized Segmentation

CVPR 2025poster

Vision Foundation Models (VFMs) excel in generalization due to large-scale pretraining, but fine-tuning them for Domain Generalized Semantic Segmentation (DGSS) while maintaining this ability remains a challenge. Existing approaches either selectively fine-tune parameters or freeze the VFMs and upda…

Cited by 0SourcePDFScholar
2025

Hybrid Contrastive Learning Decoupling Speech Emotion Recognition

ICASSP 2025accepted

Speech signals contain rich information, such as textual content, emotion, and speaker identity. To extract these features more efficiently, researchers are investigating joint training across multiple tasks, like Speech Emotion Recognition (SER) and Speaker Verification (SV), aiming to improve perf…

Cited by 0SourceScholar
2025

Predicting Spectral Information for Self-Supervised Signal Classification

IJCAI 2025

Deep learning methods have demonstrated remarkable performance across various communication signal processing tasks. However, most signal classification methods require a substantial amount of labeled samples for training, posing significant challenges in the field of communication signals, as label

Cited by 0SourcePDFScholar
2025

Pseudo-SD: Pseudo Controlled Stable Diffusion for Semi-Supervised and Cross-Domain Semantic Segmentation

ICCV 2025poster

Pseudo-labeling is a key technique of semi-supervised and cross-domian semantic segmentation, yet its efficacy is often hampered by the intrinsic noise of pseudo-labels. This study introduces Pseudo-SD, a novel framework that redefines the utilization of pseudo-label knowledge through Stable Diffusi…

2025

STEM-LTS: Integrating Semantic-Temporal Dynamics in LLM-driven Time Series Analysis

AAAI 2025technical

Time series forecasting plays a crucial role in domains such as finance, healthcare, and climate science. However, as modern time series data become increasingly complex, featuring high dimensionality, intricate spatiotemporal dependencies, and multi-scale evolutionary patterns, traditional analytic…

Cited by 0SourcePDFScholar
2024

Connectivity-Driven Pseudo-Labeling Makes Stronger Cross-Domain Segmenters

NeurIPS 2024poster

Presently, pseudo-labeling stands as a prevailing approach in cross-domain semantic segmentation, enhancing model efficacy by training with pixels assigned with reliable pseudo-labels. However, we identify two key limitations within this paradigm: (1) under relatively severe domain shifts, most sel…

Cited by 1SourcePDFScholar
2024

Disentangle Estimation of Causal Effects from Cross-Silo Data

ICASSP 2024accepted

Estimating causal effects among different events is of great importance to critical fields such as drug development. Nevertheless, the data features associated with events may be distributed across various silos and remain private within respective parties, impeding direct information exchange betwe…

Cited by 0SourceScholar
2024

Stable Neighbor Denoising for Source-free Domain Adaptive Segmentation

CVPR 2024poster

We study source-free unsupervised domain adaptation (SFUDA) for semantic segmentation which aims to adapt a source-trained model to the target domain without accessing the source data. Many works have been proposed to address this challenging problem among which uncertainty based self-training is a…

2024

VSViG: Real-time Video-based Seizure Detection via Skeleton-based Spatiotemporal ViG

ECCV 2024poster

"An accurate and efficient epileptic seizure onset detection can significantly benefit patients. Traditional diagnostic methods, primarily relying on electroencephalograms (EEGs), often result in cumbersome and non-portable solutions, making continuous patient monitoring challenging. The video-based…

2023

Learning Pseudo-Relations for Cross-domain Semantic Segmentation

ICCV 2023poster

Domain adaptive semantic segmentation aims to adapt a model trained on labeled source domain to the unlabeled target domain. Self-training shows competitive potential in this field. Existing methods along this stream mainly focus on selecting reliable predictions on target data as pseudo-labels for…

Cited by 23PDFcodeScholar
2023

Point Clouds Outlier Removal Method Based on Improved Mahalanobis and Completion

RA-L 2023

Point clouds have been regarded as a representative format for 3D visualization of real-world objects or scenes. However, point clouds acquired from depth cameras or laser scanning devices commonly contain outliers. Outlier removal performance will directly affect the downstream applications. Existi

Cited by 11SourceScholar
2023

Towards Better Stability and Adaptability: Improve Online Self-Training for Model Adaptation in Semantic Segmentation

CVPR 2023highlight

Unsupervised domain adaptation (UDA) in semantic segmentation transfers the knowledge of the source domain to the target one to improve the adaptability of the segmentation model in the target domain. The need to access labeled source data makes UDA unable to handle adaptation scenarios involving pr…

2022

Optimal Control for a Modified Bouc-Wen Model in a Magnetorheological Fluid Master Robot

RA-L 2022

Magnetorheological fluid (MRF) clutch is a kind of passive actuator with advantages, like fast response and low inertia. However, due to the complex natural property, it is difficult to model such devices accurately, and this leads to degraded performance. In this letter, a steady-state model and tr

Cited by 4SourceScholar
2021

Cascade Attention Fusion for Fine-Grained Image Captioning Based on Multi-Layer LSTM

ICASSP 2021accepted

The conventional visual attention-based image captioning approaches typically use image information to guide caption generation. Results from these models tend to be coarse and ignore the details in the image, such as objects, attributes and the distinguishing aspects of each image. In this paper, w…

Cited by 0SourceScholar
2019

AFD-Net: Aggregated Feature Difference Learning for Cross-Spectral Image Patch Matching

ICCV 2019oral

Image patch matching across different spectral domains is more challenging than in a single spectral domain. We consider the reason is twofold: 1. the weaker discriminative feature learned by conventional methods; 2. the significant appearance difference between two images domains. To tackle these p…

Cited by 38PDFScholar
2019

Better and Faster: Exponential Loss for Image Patch Matching

ICCV 2019poster

Recent studies on image patch matching are paying more attention on hard sample learning, because easy samples do not contribute much to the network optimization. They have proposed various hard negative sample mining strategies, but very few addressed this problem from the perspective of loss funct…

Cited by 31PDFScholar
2019

Language Person Search with Mutually Connected Classification Loss

ICASSP 2019accepted

In this work, we develop an effective person search algorithm with natural language descriptions. The contributions of this work mainly include two aspects. First, we design a baseline language person search framework including three basic components: a deep CNN model to extract visual features, a b…

Cited by 0SourceScholar
2019

Online Single Person Tracking for Unmanned Aerial Vehicles: Benchmark and New Baseline

ICASSP 2019accepted

Online tracking a specific person from a low-altitude unmanned aerial vehicle (UAV) is a very interesting and challenging problem to be solved. However, there exists no large-scale aerial video dataset regarding this online single person tracking (OSPT) task. To promote the study of the OSPT problem…

Cited by 0SourceScholar
2018

Real-time 'Actor-Critic' Tracking

ECCV 2018poster

In this work, we propose a novel tracking algorithm with real-time performance based on the ‘Actor-Critic’ framework. This framework consists of two major components: ‘Actor’ and ‘Critic’. The ‘Actor’ model aims to infer the optimal choice in a continuous action space, which directly makes the track…