DSSG: Dual-Stream Semantic Guidance for Source-Fully-Free Adaptation of Vision-Language Models
Weiwei Xiang, Shun Peng, Guangyi Xiao, Hao Chen, Lei Yang
Abstract
Source-Fully-Free Domain Adaptation (SFF-DA) has emerged as a strategic paradigm to adapt Vision-Language Models (VLMs) without any access to source data or task-specific source models. However, we identify a critical "Dual Semantic Drift" that hinders this process: static drift caused by the stagnation of frozen hand-crafted templates, and dynamic drift resulting from noisy, instance-level generated captions. These issues lead to a severe stability-plasticity dilemma, manifesting as semantic misalignment and class collapse. To address this, we propose DSSG (Dual-Stream Semantic Guidance), an end-to-end framework that reconciles fine-grained plasticity with global stability. Our core contribution is the Dual Semantic Guidance (DSG) module, which integrates a captioning stream for detailed domain nuances with a template stream to anchor global categorical consistency. Furthermore, a Dynamic Cross-Modal Knowledge Distillation (CMKD) module is introduced to adaptively optimize teacher-student alignment across modalities. Extensive experiments demonstrate that DSSG significantly outperforms current state-of-the-art methods across multiple benchmarks, providing a robust solution for the unsupervised transfer of foundation models. The code is available on https://github.com/mrmenand/DSSG.
BibTeX
@inproceedings{ijcai2026_dssgdualstreamse,
title = {DSSG: Dual-Stream Semantic Guidance for Source-Fully-Free Adaptation of Vision-Language Models},
author = {Weiwei Xiang and Shun Peng and Guangyi Xiao and Hao Chen and Lei Yang},
booktitle = {IJCAI 2026},
year = {2026}
}