Global Compression Commander: Plug-and-Play Inference Acceleration for High-Resolution Large Vision-Language Models
Large vision-language models (LVLMs) excel at visual understanding but face efficiency challenges due to quadratic complexity when processing long multimodal contexts. While token compression can reduce computational costs, existing approaches are designed for single-view LVLMs and fail to account f