← Search

Tokio Kajitsuka

2 accepted papers

2024

Are Transformers with One Layer Self-Attention Using Low-Rank Weight Matrices Universal Approximators?

ICLR 2024poster

Existing analyses of the expressive capacity of Transformer models have required excessively deep layers for data memorization, leading to a discrepancy with the Transformers actually used in practice. This is primarily due to the interpretation of the softmax function as an approximation of the ha…

Cited by 28SourcePDFScholar