← Search

Sifan Li

1 accepted papers

2025

SemVink: Advancing VLMs’ Semantic Understanding of Optical Illusions via Visual Global Thinking

EMNLP 2025

Vision-language models (VLMs) excel in semantic tasks but falter at a core human capability: detecting hidden content in optical illusions or AI-generated images through perceptual adjustments like zooming. We introduce HC-Bench, a benchmark of 112 images with hidden texts, objects, and illusions, r

Cited by 0SourcePDFScholar