← Search

Chenggong Zhao

1 accepted papers

2026

Reading Images Like Texts: Sequential Image Understanding in Vision-Language Models

ICLR 2026poster

Vision-Language Models (VLMs) have demonstrated remarkable performance across a variety of real-world tasks. However, existing VLMs typically process visual information by serializing images, a method that diverges significantly from the parallel nature of human vision. Moreover, their opaque intern…

Cited by 0SourcecodeScholar