← Search

Shu Zou

5 accepted papers

2026

All Roads Lead to Rome: Incentivizing Divergent Thinking in Vision-Language Models

CVPR 2026

Recent studies have demonstrated that Reinforcement Learning (RL), notably Group Relative Policy Optimization (GRPO), can intrinsically elicit and enhance the reasoning capabilities of Vision-Language Models (VLMs). However, despite the promise, the underlying mechanisms that drive the effectiveness

Cited by 0SourcecodeScholar
2026

More Thought, Less Accuracy? On the Dual Nature of Reasoning in Vision-Language Models

ICLR 2026poster

Reasoning has emerged as a pivotal capability in Large Language Models (LLMs). Through Reinforcement Learning (RL), typically Group Relative Policy Optimization (GRPO), these models are able to solve complex tasks such as mathematics and code generation. Building on these advances, recent research h…

Cited by 0SourcecodeScholar
2025

Black Sheep in the Herd: Playing with Spuriously Correlated Attributes for Vision-Language Recognition

ICLR 2025poster

Few-shot adaptation for Vision-Language Models (VLMs) presents a dilemma: balancing in-distribution accuracy with out-of-distribution generalization. Recent research has utilized low-level concepts such as visual attributes to enhance generalization. However, this study reveals that VLMs overly rely…

Cited by 0SourcePDFScholar
2025

Identifying and Mitigating Position Bias of Multi-image Vision-Language Models

CVPR 2025poster

The evolution of Large Vision-Language Models (LVLMs) has progressed from single-image understanding to multi-image reasoning. Despite this advancement, our findings indicate that LVLMs struggle to robustly utilize information across multiple images, with predictions significantly affected by the al…