HiDivDrop: Vision Token Reduction in MLLMs via Late Injection and Differentiable Top-K
The computational cost of Multimodal Large Language Models (MLLMs), driven by the quadratic complexity of processing vision tokens, remains a significant barrier to their widespread adoption. While progressive vision token pruning is a promising solution, we find that its full potential has been unr…