2026
Improving Visual Token Reduction via Rectifying Distortions for Efficient Multimodal LLM Inference
ICML 2026poster
Recent advancements in Multimodal Large Language Models (MLLMs) have achieved remarkable success in vision-language tasks, yet the quadratic computational complexity arising from the vast number of visual tokens creates significant memory and latency bottlenecks. While visual token reduction (VTR) s…