2025
Continuous-Time Attention: PDE-Guided Mechanisms for Long-Sequence Transformers
EMNLP 2025
We present Continuous-Time Attention, a novel framework that infuses partial differential equations (PDEs) into the Transformer’s attention mechanism to better handle long sequences. Instead of relying on a static attention matrix, we allow attention weights to evolve along a pseudo-time dimension g