2023
Continual Transformers: Redundancy-Free Attention for Online Inference
ICLR 2023poster
Transformers in their common form are inherently limited to operate on whole token sequences rather than on one token at a time. Consequently, their use during online inference on time-series data entails considerable redundancy due to the overlap in successive token sequences. In this work, we prop…