iLLaVA: An Image is Worth Fewer Than 1/3 Input Tokens in Large Multimodal Models
Recent methods have made notable progress in accelerating Large Vision-Language Models (LVLMs) by exploiting the inherent redundancy in visual inputs. Most existing approaches, however, focus narrowly on reducing image tokens before or within the Large Language Model (LLM) stage to lower computation…