2026
Inverse Optimal Transport for Efficient Adaptation of Vision-Language Models
AAAI 2026technical
Vision–language models (VLMs) such as CLIP have unlocked powerful zero-shot transfer, yet efficient adaptation to downstream tasks remains challenging. Existing methods often depend on graph structures and dataset-specific tuning, making them sensitive to modality gaps and computationally costly at