ICML 2026poster0 citations

Intrinsic Gradient Suppression for Label-Noise Prompt Tuning in Vision–Language Models

Jia-yu Li, Jiaxin Qi, Sheng Zhou, Jianqiang Huang, Xian-Sheng Hua

Abstract

Contrastive vision-language models like CLIP exhibit remarkable zero-shot generalization. However, prompt tuning remains highly sensitive to label noise, as mislabeled samples generate disproportionately large gradients that can overwhelm pre-trained priors. We argue that because CLIP already provides a near-optimal initialization, adaptation should be inherently conservative, particularly against the extreme gradient updates common in noisy settings. To this end, we propose Double-Softmax Prompt Tuning (DSPT), a hyperparameter-free method for intrinsic gradient suppression. By applying a sequential probabilistic normalization, DSPT induces a self-adaptive saturation zone that suppresses gradients from high-error noisy samples while maintaining informative updates. We also provide both theoretical analysis and empirical evidence about how this mechanism achieves adaptive suppression. This design transforms ``gradient vanishing'', traditionally a training bottleneck, into a principled noise-filtering shield for label-noise prompt tuning. Extensive experiments confirm that this simple, drop-in design achieves state-of-the-art robustness across various noisy benchmarks, outperforming methods with complex architectures and handcrafted hyperparameters.

OptimizationTheoryRobustnessVisionMultimodalBenchmark
BibTeX
@inproceedings{
li2026intrinsic,
title={Intrinsic Gradient Suppression for Label-Noise Prompt Tuning in Vision{\textendash}Language Models},
author={Jia-yu Li and Jiaxin Qi and Sheng Zhou and Jianqiang Huang and Xian-Sheng Hua},
booktitle={Forty-third International Conference on Machine Learning},
year={2026},
url={https://openreview.net/forum?id=9W17xHFtdQ}
}