IJCAI 2023poster21 citations

Black-box Prompt Tuning for Vision-Language Model as a Service

Lang Yu, Qin Chen, Jiaju Lin, Liang He

Abstract

In the scenario of Model-as-a-Service (MaaS), pre-trained models are usually released as inference APIs. Users are allowed to query those models with manually crafted prompts. Without accessing the network structure and gradient information, it's tricky to perform continuous prompt tuning on MaaS, especially for vision-language models (VLMs) considering cross-modal interaction. In this paper, we propose a black-box prompt tuning framework for VLMs to learn task-relevant prompts without back-propagation. In particular, the vision and language prompts are jointly optimized in the intrinsic parameter subspace with various evolution strategies. Different prompt variants are also explored to enhance the cross-model interaction. Experimental results show that our proposed black-box prompt tuning framework outperforms both hand-crafted prompt engineering and gradient-based prompt learning methods, which serves as evidence of its capability to train task-relevant prompts in a derivative-free manner.

Computer Vision: CV: Vision and languageMachine Learning: ML: Evolutionary learningMachine Learning: ML: Multi-modal learning
BibTeX
@inproceedings{ijcai2023p187,
  title     = {Black-box Prompt Tuning for Vision-Language Model as a Service},
  author    = {Yu, Lang and Chen, Qin and Lin, Jiaju and He, Liang},
  booktitle = {Proceedings of the Thirty-Second International Joint Conference on
               Artificial Intelligence, {IJCAI-23}},
  publisher = {International Joint Conferences on Artificial Intelligence Organization},
  editor    = {Edith Elkind},
  pages     = {1686--1694},
  year      = {2023},
  month     = {8},
  note      = {Main Track},
  doi       = {10.24963/ijcai.2023/187},
  url       = {https://doi.org/10.24963/ijcai.2023/187},
}
Black-box Prompt Tuning for Vision-Language Model as a Service · IJCAI 2023