An Empirical Study of Attention and Diversity for Adaptive Visual Token Pruning in Large Vision-Language Models
Large Vision-Language Models (LVLMs) have adopted visual token pruning strategies to mitigate substantial computational overhead incurred by extensive visual token sequences. While prior works primarily focus on either attention-based or diversity-based pruning methods, in-depth analysis of these a…