← Search

Xingzhou Pang

1 accepted papers

2026

Unveiling the Visual Counting Bottleneck in Vision-Language Models

ICML 2026poster

While Large Vision-Language Models (VLMs) excel at interpolation, they suffer catastrophic failures in systematic generalization, most notably in visual counting beyond training distributions. In this work, we investigate this extrapolation bottleneck by deconstructing visual counting into three cog…

Cited by 0SourceScholar