ICML 2025poster1 citations

On Zero-Initialized Attention: Optimal Prompt and Gating Factor Estimation

Nghiem Tuong Diep, Huy Nguyen, Chau Nguyen, Minh Le, Duy Minh Ho Nguyen, Daniel Sonntag, Mathias Niepert, Nhat Ho

Abstract

LLaMA-Adapter has recently emerged as an efficient fine-tuning technique for LLaMA models, leveraging zero-initialized attention to stabilize training and enhance performance. However, despite its empirical success, the theoretical foundations of zero-initialized attention remain largely unexplored. In this paper, we provide a rigorous theoretical analysis, establishing a connection between zero-initialized attention and mixture-of-expert models. We prove that both linear and non-linear prompts, along with gating functions, can be optimally estimated, with non-linear prompts offering greater flexibility for future applications. Empirically, we validate our findings on the open LLM benchmarks, demonstrating that non-linear prompts outperform linear ones. Notably, even with limited training data, both prompt types consistently surpass vanilla attention, highlighting the robustness and adaptability of zero-initialized attention.

Large Language Model (LLM)TheoryInstruction TuningMixture of Experts
BibTeX
@inproceedings{
diep2025on,
title={On Zero-Initialized Attention: Optimal Prompt and Gating Factor Estimation},
author={Nghiem Tuong Diep and Huy Nguyen and Chau Nguyen and Minh Le and Duy Minh Ho Nguyen and Daniel Sonntag and Mathias Niepert and Nhat Ho},
booktitle={Forty-second International Conference on Machine Learning},
year={2025},
url={https://openreview.net/forum?id=auc4Ng6h4R}
}
On Zero-Initialized Attention: Optimal Prompt and Gating Factor Estimation · ICML 2025