2025
Text or Pixels? Evaluating Efficiency and Understanding of LLMs with Visual Text Inputs
EMNLP 2025
Large language models (LLMs) and their multimodal variants can now process visual inputs, including images of text. This raises an intriguing question: Can we compress textual inputs by feeding them as images to reduce token usage while preserving performance?In this paper, we show that *visual text