2026
CATP: Contextually Adaptive Token Pruning for Efficient and Enhanced Multimodal In-Context Learning
AAAI 2026technical
Modern large vision-language models (LVLMs) convert each input image into a large set of tokens that far outnumber the text tokens. Although this improves visual perception, it also introduces severe image token redundancy. Because image tokens contain sparse information, many contribute little to r