2025
SemVink: Advancing VLMs’ Semantic Understanding of Optical Illusions via Visual Global Thinking
EMNLP 2025
Vision-language models (VLMs) excel in semantic tasks but falter at a core human capability: detecting hidden content in optical illusions or AI-generated images through perceptual adjustments like zooming. We introduce HC-Bench, a benchmark of 112 images with hidden texts, objects, and illusions, r