EMNLP 2023long main0 citations

APrompt: Attention Prompt Tuning for Efficient Adaptation of Pre-trained Language Models

Qifan Wang, Yuning Mao, Jingang Wang, Hanchao Yu, Shaoliang Nie, Sinong Wang, Fuli Feng, Lifu Huang

Abstract

With the continuous growth of large language models, the process of fine-tuning these models for new tasks has become increasingly parameter-intensive. Prompt tuning, a method that involves tuning a small set of soft prompts, has emerged as an effective and efficient approach for adapting large pre-trained language models. However, most existing prompt tuning approaches only introduce prompts at the input layer, limiting their performance and leaving large rooms for improvement. In this work, we propose a novel Attention Prompt tuning method, namely APrompt, for efficient adaptation of pre-trained language models. We first demonstrate that existing prompt tuning can be considered as a special case of attention prompt tuning. We then formally introduce APrompt, which incorporates query, key, and value prompts into the attention layer to guide the attention computation during fine-tuning. Experimental results on the SuperGLUE benchmark consistently demonstrate that our proposed approach outperforms state-of-the-art baselines and full fine-tuning method with pre-trained models at different scales. In addition, a comprehensive set of ablation studies validate the effectiveness of the prompt design, as well as the efficiency of our approach.

Prompt TuningParameter Efficient LearningAttention Prompt
BibTeX
@inproceedings{
wang2023aprompt,
title={{AP}rompt: Attention Prompt Tuning for Efficient Adaptation of Pre-trained Language Models},
author={Qifan Wang and Yuning Mao and Jingang Wang and Hanchao Yu and Shaoliang Nie and Sinong Wang and Fuli Feng and Lifu Huang and Xiaojun Quan and Zenglin Xu and Dongfang Liu},
booktitle={The 2023 Conference on Empirical Methods in Natural Language Processing},
year={2023},
url={https://openreview.net/forum?id=RSuN6p3wXR}
}
APrompt: Attention Prompt Tuning for Efficient Adaptation of Pre-trained Language Models · EMNLP 2023