Grid-Attention: Enhancing Computational Efficiency of Large Vision Models without Fine-Tuning
"Recently, transformer-based large vision models, , the Segment Anything Model (SAM) and Stable Diffusion (SD), have achieved remarkable success in the computer vision field. However, the quartic complexity within the transformer’s Multi-Head Attention (MHA) leads to substantial computational costs…