2024
FlexAttention for Efficient High-Resolution Vision-Language Models
ECCV 2024poster
"Current high-resolution vision-language models encode images as high-resolution image tokens and exhaustively take all these tokens to compute attention, which significantly increases the computational cost. To address this problem, we propose , a flexible attention mechanism for efficient high-res…