2026
Efficient Multimodal Large Language Model via Dynamic KV Cache Quantization
AAAI 2026technical
Multimodal large language models (LMMs) have demonstrated remarkable capabilities across diverse vision-language tasks, including image captioning, visual question answering, and text-image retrieval. However, their computational complexity and memory footprint, particularly in the key-value (KV) ca