2026
On Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMs
ICML 2026poster
Reinforcement learning (RL) fine-tuning is now widely used to improve LLM reasoning, and recent work has begun extending it to vision-language models (VLMs). While RL-tuned VLMs can improve visual reasoning benchmark performance, they can still suffer from weak visual grounding, hallucinations, and …