Efficient and Exact Global Attention on Latent Summaries for Knowledge Graph Reasoning
Capturing global context through attention is essential for reasoning over knowledge graphs, especially when relevant entities are distant or disconnected. To scale attention to large graphs, recent methods replace Softmax with kernel feature mappings, reducing computational complexity to linear in