2026
Remodeling Semantic Relationships in Vision-Language Fine-Tuning
AAAI 2026technical
Vision-language fine-tuning has emerged as an efficient paradigm for constructing multimodal foundation models. While textual context often highlights semantic relationships within an image, existing fine-tuning methods typically overlook this information when aligning vision and language, thus lead