2025
LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
ICCV 2025poster
Large language models have demonstrated substantial advancements in reasoning capabilities. However, current Vision-Language Models (VLMs) often struggle to perform systematic and structured reasoning, especially when handling complex visual question-answering tasks. In this work, we introduce LLaVA…