Filter, Correlate, Compress: Training-Free Token Reduction for MLLM Acceleration
The quadratic complexity of Multimodal Large Language Models (MLLMs) with respect to context length poses significant computational and memory challenges, hindering their real-world deployment. In the paper, we devise a