← Search

Huiyin Xue

3 accepted papers

2023

Pit One Against Many: Leveraging Attention-head Embeddings for Parameter-efficient Multi-head Attention

EMNLP 2023long findings

Scaling pre-trained language models has resulted in large performance gains in various natural language processing tasks but comes with a large cost in memory requirements. Inspired by the position embeddings in transformers, we aim to simplify and reduce the memory footprint of the multi-head atten…

Cited by 0SourceScholar