LLaVA-FA: Learning Fourier Approximation for Compressing Large Multimodal Models
Large multimodal models (LMMs) have achieved impressive performance on various vision-language tasks, but their substantial computational and memory costs hinder their practical deployment. Existing compression methods often decouple low-rank decomposition and quantization, leading to compounded rec…