2026
Instruction-Guided Cross-Modal Clustering for Training-Free Visual Token Pruning in Vision-Language Models
AAAI 2026technical
Large vision-language models (LVLMs) have demonstrated remarkable capabilities in understanding multimodal data such as images and text. However, the number of visual tokens in these models often far exceeds that of textual tokens, resulting in substantial redundancy and high inference costs. Existi