← Search

Yuji Yamamoto

1 accepted papers

2023

Absolute Position Embedding Learns Sinusoid-like Waves for Attention Based on Relative Position

EMNLP 2023long main

Attention weight is a clue to interpret how a Transformer-based model makes an inference. In some attention heads, the attention focuses on the neighbors of each token. This allows the output vector of each token to depend on the surrounding tokens and contributes to make the inference context-depen…

Cited by 0SourceScholar