← Search

Yunya Zhang

1 accepted papers

2026

Instruction-Guided Cross-Modal Clustering for Training-Free Visual Token Pruning in Vision-Language Models

AAAI 2026technical

Large vision-language models (LVLMs) have demonstrated remarkable capabilities in understanding multimodal data such as images and text. However, the number of visual tokens in these models often far exceeds that of textual tokens, resulting in substantial redundancy and high inference costs. Existi

Cited by 0SourcePDFScholar