Hi-Lo Prune: Look at What You'll Lose before Pruning with Hierarchical Token Selection
Multimodal Large Language Models (MLLMs) have achieved remarkable progress in vision-language understanding, yet processing long visual token sequences remains computationally expensive. Existing approaches mitigate this cost by reducing image tokens, either by discarding them after the visual encod