AAAI 2026technical0 citations

Global Compression Commander: Plug-and-Play Inference Acceleration for High-Resolution Large Vision-Language Models

Xuyang Liu, Ziming Wang, Junjie Chen, Yuhang Han, Yingyao Wang, Jiale Yuan, Jun Song, Siteng Huang

Abstract

Large vision-language models (LVLMs) excel at visual understanding but face efficiency challenges due to quadratic complexity when processing long multimodal contexts. While token compression can reduce computational costs, existing approaches are designed for single-view LVLMs and fail to account for the unique multi-view characteristics of high-resolution LVLMs that use dynamic cropping. Current methods treat all tokens uniformly, yet our analysis shows that global thumbnails can naturally guide the compression of local crops by providing holistic context for evaluating informativeness. In this paper, we first analyze the dynamic cropping strategy, revealing both the complementary relationship between thumbnails and crops and the distinct characteristics across different crops. Based on these insights, we propose

BibTeX
@inproceedings{aaai2026_globalcompressio,
  title = {Global Compression Commander: Plug-and-Play Inference Acceleration for High-Resolution Large Vision-Language Models},
  author = {Xuyang Liu and Ziming Wang and Junjie Chen and Yuhang Han and Yingyao Wang and Jiale Yuan and Jun Song and Siteng Huang and Honggang Chen},
  booktitle = {AAAI 2026},
  year = {2026}
}