2026
Prompt-Robust Vision-Language Models via Meta-Finetuning
ICLR 2026poster
Vision-language models (VLMs) have demonstrated remarkable generalization across diverse tasks by leveraging large-scale image-text pretraining. However, their performance is notoriously unstable under variations in natural language prompts, posing a considerable challenge for reliable real-world de…