2025
Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs
EMNLP 2025
Prompt sensitivity, referring to the phenomenon where paraphrasing (that is, repeating something written or spoken using different words) leads to significant changes in large language model performance, has been widely accepted as a core limitation of large language models. In this work, we revisit