2026
PAR: Prompt-Aware Token Reduction Method for Efficient Large Multimodal Models
ICASSP 2026poster
Multimodal large language models (MLLMs) demonstrate strong performance across visual tasks, but their efficiency is hindered by significant computational and memory demands from processing long contexts in multimodal inputs. To address this, we introduce PAR (Prompt-Aware Token Reduction), a novel…