2025
Inference Optimal VLMs Need Fewer Visual Tokens and More Parameters
ICLR 2025poster
Vision Language Models (VLMs) have demonstrated strong capabilities across various visual understanding and reasoning tasks, driven by incorporating image representations into the token inputs of Large Language Models (LLMs). However, their real-world deployment is often constrained by high latency…