← Search

Difan Deng

2 accepted papers

2026

Neural Attention Search Linear: Towards Adaptive Token-Level Hybrid Attention Models

ICML 2026poster

The quadratic computational complexity of softmax transformers has become a bottleneck in long-context scenarios. In contrast, linear attention model families provide a promising direction towards a more efficient sequential model. These linear attention models compress past $KV$ values into a singl…

Cited by 0SourceScholar
2025

Neural Attention Search

NeurIPS 2025poster

We present Neural Attention Search (NAtS), an end-to-end learnable sparse transformer that automatically evaluates the importance of each token within a sequence and determines if the corresponding token can be dropped after several steps. To this end, we design a search space that contains three to…

Cited by 0SourceScholar