← Search

Yaguang Song

5 accepted papers

2026

Dance Across Shifts: Forward-Facilitation Continual Test-Time Adaptation through Dynamic Style Bridging

CVPR 2026

Continual Test-Time Adaptation (CTTA) aims to empower perception systems to handle dynamic distribution shifts encountered after deployment. Existing methods predominantly follow a backward-alignment paradigm, which rigidly aligns incoming data with supervisory surrogates derived from the source dom

Cited by 0SourcecodeScholar
2026

DuetMerging: Synergizing Dynamic and Static Strategies for Mitigating Task Interference in Model Merging

CVPR 2026

Model merging offers a promising paradigm for consolidating multiple expert models into a single multitask architecture. However, its effectiveness is often hindered by task interference, where conflicting parameter updates from different tasks degrade performance. While dynamic, Mixture-of-Experts

Cited by 0SourceScholar
2025

Pilot: Building the Federated Multimodal Instruction Tuning Framework

AAAI 2025technical

In this paper, we explore a novel federated multimodal instruction tuning task(FedMIT), which is significant for collaboratively fine-tuning MLLMs on different types of multimodal instruction data on distributed devices. To solve the new task, we propose a federated multimodal instruction tuning fra…

Cited by 1SourcePDFScholar
2024

Libra: Building Decoupled Vision System on Large Language Models

ICML 2024poster

In this work, we introduce **Libra**, a prototype model with a decoupled vision system on a large language model (LLM). The decoupled vision system decouples inner-modal modeling and cross-modal interaction, yielding unique visual information modeling and effective cross-modal comprehension. Libra i…

2024

Modality-Collaborative Test-Time Adaptation for Action Recognition

CVPR 2024poster

Video-based Unsupervised Domain Adaptation (VUDA) method improves the generalization of the video model enabling it to be applied to action recognition tasks in different environments. However these methods require continuous access to source data during the adaptation process which are impractical…

Cited by 6SourcePDFScholar