ICML 2022spotlight4 citations

Ripple Attention for Visual Perception with Sub-quadratic Complexity

Lin Zheng, Huijie Pan, Lingpeng Kong

Abstract

Transformer architectures are now central to sequence modeling tasks. At its heart is the attention mechanism, which enables effective modeling of long-term dependencies in a sequence. Recently, transformers have been successfully applied in the computer vision domain, where 2D images are first segmented into patches and then treated as 1D sequences. Such linearization, however, impairs the notion of spatial locality in images, which bears important visual clues. To bridge the gap, we propose

BibTeX
@InProceedings{pmlr-v162-zheng22a,
  title = 	 {Ripple Attention for Visual Perception with Sub-quadratic Complexity},
  author =       {Zheng, Lin and Pan, Huijie and Kong, Lingpeng},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {26993--27010},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/zheng22a/zheng22a.pdf},
  url = 	 {https://proceedings.mlr.press/v162/zheng22a.html},
  abstract = 	 {Transformer architectures are now central to sequence modeling tasks. At its heart is the attention mechanism, which enables effective modeling of long-term dependencies in a sequence. Recently, transformers have been successfully applied in the computer vision domain, where 2D images are first segmented into patches and then treated as 1D sequences. Such linearization, however, impairs the notion of spatial locality in images, which bears important visual clues. To bridge the gap, we propose