2025
FlashSloth : Lightning Multimodal Large Language Models via Embedded Visual Compression
CVPR 2025poster
Despite a big leap forward in capability, multimodal large language models (MLLMs) tend to behave like a sloth in practical use, i.e., slow response and large latency. Recent efforts are devoted to building tiny MLLMs for better efficiency, but the plethora of visual tokens still used limit their ac…