2025
Cache-of-Thought: Master-Apprentice Framework for Cost-Effective Vision Language Model Reasoning
EMNLP 2025
Vision Language Models (VLMs) have achieved remarkable success in a wide range of vision applications of increasing complexity and scales, yet choosing the right VLM model size involves a trade-off between response quality and cost. While smaller VLMs are cheaper to run, they typically produce respo