← Search

Amir Ajorlou

2 accepted papers

2024

On the Role of Attention Masks and LayerNorm in Transformers

NeurIPS 2024poster

Self-attention is the key mechanism of transformers, which are the essential building blocks of modern foundation models. Recent studies have shown that pure self-attention suffers from an increasing degree of rank collapse as depth increases, limiting model expressivity and further utilization of m…

Cited by 12SourcePDFScholar
2023

Demystifying Oversmoothing in Attention-Based Graph Neural Networks

NeurIPS 2023spotlight

Oversmoothing in Graph Neural Networks (GNNs) refers to the phenomenon where increasing network depth leads to homogeneous node representations. While previous work has established that Graph Convolutional Networks (GCNs) exponentially lose expressive power, it remains controversial whether the grap…

Cited by 57SourcePDFScholar