2025
Textural or Textual: How Vision-Language Models Read Text in Images
ICML 2025poster
Typographic attacks are often attributed to the ability of multimodal pre-trained models to fuse textual semantics into visual representations, yet the mechanisms and locus of such interference remain unclear. We examine whether such models genuinely encode textual semantics or primarily rely on tex…