AAAI 2024technical7 citations

Cached Transformers: Improving Transformers with Differentiable Memory Cachde

Zhaoyang Zhang, Wenqi Shao, Yixiao Ge, Xiaogang Wang, Jinwei Gu, Ping Luo

Abstract

This work introduces a new Transformer model called Cached Transformer, which uses Gated Recurrent Cached (GRC) attention to extend the self-attention mechanism with a differentiable memory cache of tokens. GRC attention enables attending to both past and current tokens, increasing the receptive field of attention and allowing for exploring long-range dependencies. By utilizing a recurrent gating unit to continuously update the cache, our model achieves significant advancements in \textbf{six} language and vision tasks, including language modeling, machine translation, ListOPs, image classification, object detection, and instance segmentation. Furthermore, our approach surpasses previous memory-based techniques in tasks such as language modeling and displays the ability to be applied to a broader range of situations.

BibTeX
@article{Zhang_Shao_Ge_Wang_Gu_Luo_2024, title={Cached Transformers: Improving Transformers with Differentiable Memory Cachde}, volume={38}, url={https://ojs.aaai.org/index.php/AAAI/article/view/29636}, DOI={10.1609/aaai.v38i15.29636}, abstractNote={This work introduces a new Transformer model called Cached Transformer, which uses Gated Recurrent Cached (GRC) attention to extend the self-attention mechanism with a differentiable memory cache of tokens. GRC attention enables attending to both past and current tokens, increasing the receptive field of attention and allowing for exploring long-range dependencies. By utilizing a recurrent gating unit to continuously update the cache, our model achieves significant advancements in \textbf{six} language and vision tasks, including language modeling, machine translation, ListOPs, image classification, object detection, and instance segmentation. Furthermore, our approach surpasses previous memory-based techniques in tasks such as language modeling and displays the ability to be applied to a broader range of situations.}, number={15}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, author={Zhang, Zhaoyang and Shao, Wenqi and Ge, Yixiao and Wang, Xiaogang and Gu, Jinwei and Luo, Ping}, year={2024}, month={Mar.}, pages={16935-16943} }
Cached Transformers: Improving Transformers with Differentiable Memory Cachde · AAAI 2024