2023
Words, Subwords, and Morphemes: What Really Matters in the Surprisal-Reading Time Relationship?
EMNLP 2023short findings
An important assumption that comes with using LLMs on psycholinguistic data has gone unverified. LLM-based predictions are based on subword tokenization, not decomposition of words into morphemes. Does that matter? We carefully test this by comparing surprisal estimates using orthographic, morpholog…