VISion On Request: Enhanced VLLM efficiency with sparse, dynamically selected, vision-language interactions
Existing approaches for improving the efficiency of Large Vision-Language Models (LVLMs) are largely based on the concept of visual token reduction. This approach, however, creates an information bottleneck that impairs performance, especially on challenging tasks that require fine-grained understan